Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

F2Bench



F2Bench is an open-ended benchmark for evaluating fairness in large language models (LLMs) that also takes factual accuracy into account. Most existing fairness benchmarks use multiple-choice formats, which differ from how people actually use LLMs, and models can score well on them by guessing. Many benchmarks also treat any difference between demographic groups as bias, even when real statistics support it, such as the higher proportion of women in nursing. F²Bench addresses both limitations. It was developed by researchers at Inner Mongolia University and Xiamen University and presented at the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP).

The benchmark contains 2 568 test instances across ten demographic categories: gender, race, religion, age, socioeconomic status, education, LGBTQ+, nationality, health and appearance. Many instances combine two categories to capture intersectional bias. Each instance contains a short scenario with two people from different groups, whose behaviours are described without saying which person is which. It also includes a verified fact, drawn from sources such as the US Census Bureau and the International Labour Organization.

Models are tested on two tasks. In the first, a six-round conversation gradually introduces stereotypical suggestions, to test whether the model eventually accepts a biased conclusion. In the second, the model describes a scenario using the given fact. It is scored on both avoiding stereotypes and reflecting the fact accurately, and these are combined into a single balance score. Responses are assessed by GPT-4 as a judge, which agreed with human reviewers in 97% of cases.

Auto-discovered on 2026-07-22 by OECD Catalogue Automation

Use Cases

There is no use cases for this tool yet.

Would you like to submit a use case for this tool?

If you have used this tool, we would love to know more about your experience.

Add use case
Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.