These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.
LangFair
LangFair is an open-source Python library, built by a team at CVS Health, for testing large language models (LLMs) for bias and fairness problems before or after they're deployed. Instead of relying on generic, one-size-fits-all benchmarks, it is designed around the idea that bias risk depends heavily on how an LLM is actually being used.
The model's responses to these prompts are collected and then assessed using metrics for toxicity, stereotyping, counterfactual fairness (whether outputs change based on protected attributes such as race or gender), and allocational harms in classification or recommendation tasks. The library also supports adversarial testing, in which prompts are designed to surface worst-case model behaviour. A decision framework included in the tool guides users in selecting which metrics are applicable to their task type and prompt characteristics. Because the evaluation relies only on model outputs rather than internal model access, it can be applied by developers, auditors, or governance bodies working with models they do not directly control. The methodology is documented in a peer-reviewed paper published in the Journal of Open Source Software, and the code is maintained publicly on GitHub.
Auto-discovered on 2026-07-11 by OECD Catalogue Automation
About the tool
You can click on the links to see the associated tools
Tool type(s):
Objective(s):
Purpose(s):
Target sector(s):
Lifecycle stage(s):
Type of approach:
Maturity:
Usage rights:
License:
Target users:
Risk management stage(s):
Github stars:
- 262
Github forks:
- 47
Use Cases
Would you like to submit a use case for this tool?
If you have used this tool, we would love to know more about your experience.
Add use case




























