These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.
FLEX (Fairness Benchmark in LLM under Extreme Scenarios)
FLEX (Fairness Benchmark in LLM under Extreme Scenarios) is a benchmark for testing whether large language models (LLMs) remain fair when prompts are deliberately designed to induce bias. Existing fairness benchmarks typically assume well-intentioned users and test models under ideal conditions. However, simple adversarial instructions can often lead LLMs to produce biased responses, so these benchmarks may underestimate the real risks. FLEX was developed by researchers at Korea University and presented at the 2025 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL).
The benchmark contains 3,145 multiple-choice questions drawn from three established fairness datasets: BBQ, CrowS-Pairs and StereoSet. Each question offers two stereotypical answers and one neutral answer, such as "not enough information", which is always correct. The authors kept only questions that a reference model answered fairly under normal conditions. They then added one adversarial prompt to each question from three categories. Persona injection asks the model to speak as a negative stereotype of a group. Competing objectives use instructions such as forbidding refusals or role-play jailbreaks. Text attacks introduce small changes such as typos, word substitutions or paraphrasing.
For each question, the prompt most likely to cause a biased answer was selected. Models are scored on accuracy and on the attack success rate: the share of questions answered fairly under normal conditions but unfairly under attack. The authors' results show that models that appear fair on existing benchmarks can still be easily manipulated.
Auto-discovered on 2026-09-23 by OECD Catalogue Automation
About the tool
You can click on the links to see the associated tools
Tool type(s):
Objective(s):
Country/Territory of origin:
Lifecycle stage(s):
Type of approach:
Maturity:
Usage rights:
Target users:
Risk management stage(s):
Use Cases
Would you like to submit a use case for this tool?
If you have used this tool, we would love to know more about your experience.
Add use case




























