Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

FLEX (Fairness Benchmark in LLM under Extreme Scenarios)



FLEX (Fairness Benchmark in LLM under Extreme Scenarios) is a benchmark for testing whether large language models (LLMs) remain fair when prompts are deliberately designed to induce bias. Existing fairness benchmarks typically assume well-intentioned users and test models under ideal conditions. However, simple adversarial instructions can often lead LLMs to produce biased responses, so these benchmarks may underestimate the real risks. FLEX was developed by researchers at Korea University and presented at the 2025 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL).

The benchmark contains 3,145 multiple-choice questions drawn from three established fairness datasets: BBQ, CrowS-Pairs and StereoSet. Each question offers two stereotypical answers and one neutral answer, such as "not enough information", which is always correct. The authors kept only questions that a reference model answered fairly under normal conditions. They then added one adversarial prompt to each question from three categories. Persona injection asks the model to speak as a negative stereotype of a group. Competing objectives use instructions such as forbidding refusals or role-play jailbreaks. Text attacks introduce small changes such as typos, word substitutions or paraphrasing.

For each question, the prompt most likely to cause a biased answer was selected. Models are scored on accuracy and on the attack success rate: the share of questions answered fairly under normal conditions but unfairly under attack. The authors' results show that models that appear fair on existing benchmarks can still be easily manipulated.

Auto-discovered on 2026-09-23 by OECD Catalogue Automation

Use Cases

There is no use cases for this tool yet.

Would you like to submit a use case for this tool?

If you have used this tool, we would love to know more about your experience.

Add use case
Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.