These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.
CEB (Compositional Evaluation Benchmark for Fairness)
CEB (Compositional Evaluation Benchmark) is a benchmark for evaluating bias in large language models (LLMs). Many datasets exist for measuring LLM bias, but each tends to focus on one type of bias and uses its own metrics, which makes results hard to compare across datasets and models. CEB addresses this by bringing existing datasets together under a single framework with consistent metrics, and by creating new datasets to fill gaps. It was developed by researchers led by the University of Virginia.
CEB is built on a compositional taxonomy that describes each dataset along three dimensions. The first is bias type: stereotyping or toxicity. The second is social group: age, gender, race or religion. The third is the task. Direct tasks ask a model to recognise biased text or select an unbiased answer. Indirect tasks ask it to continue a text, reply in a conversation or make a classification decision, and then examine the output for bias. Combining these dimensions produces a wide range of test configurations. In total, the benchmark contains 11,004 samples.
The benchmark adapts established datasets, including BBQ, HolisticBias, Adult, Credit and Jigsaw. Where no suitable data existed, the authors used GPT-4 to generate new samples. Direct tasks are scored on accuracy, and generated text is scored for bias or toxicity by GPT-4 or the Perspective API. Classification tasks use group fairness metrics such as demographic parity and equalised odds.
Auto-discovered on 2026-09-23 by OECD Catalogue Automation
About the tool
You can click on the links to see the associated tools
Objective(s):
Purpose(s):
Country/Territory of origin:
Lifecycle stage(s):
Type of approach:
Usage rights:
Target users:
Benefits:
Risk management stage(s):
Use Cases
Would you like to submit a use case for this tool?
If you have used this tool, we would love to know more about your experience.
Add use case




























