These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.
F2Bench
F2Bench is an open-ended benchmark for evaluating fairness in large language models (LLMs) that also takes factual accuracy into account. Most existing fairness benchmarks use multiple-choice formats, which differ from how people actually use LLMs, and models can score well on them by guessing. Many benchmarks also treat any difference between demographic groups as bias, even when real statistics support it, such as the higher proportion of women in nursing. F²Bench addresses both limitations. It was developed by researchers at Inner Mongolia University and Xiamen University and presented at the 2025 Conference on Empirical Methods in Natural Language Processing (EMNLP).
The benchmark contains 2 568 test instances across ten demographic categories: gender, race, religion, age, socioeconomic status, education, LGBTQ+, nationality, health and appearance. Many instances combine two categories to capture intersectional bias. Each instance contains a short scenario with two people from different groups, whose behaviours are described without saying which person is which. It also includes a verified fact, drawn from sources such as the US Census Bureau and the International Labour Organization.
Models are tested on two tasks. In the first, a six-round conversation gradually introduces stereotypical suggestions, to test whether the model eventually accepts a biased conclusion. In the second, the model describes a scenario using the given fact. It is scored on both avoiding stereotypes and reflecting the fact accurately, and these are combined into a single balance score. Responses are assessed by GPT-4 as a judge, which agreed with human reviewers in 97% of cases.
Auto-discovered on 2026-07-22 by OECD Catalogue Automation
About the tool
You can click on the links to see the associated tools
Tool type(s):
Objective(s):
Purpose(s):
Lifecycle stage(s):
Type of approach:
Maturity:
Usage rights:
Target users:
Risk management stage(s):
Technology platforms:
Github stars:
- 2
Use Cases
Would you like to submit a use case for this tool?
If you have used this tool, we would love to know more about your experience.
Add use case




























