Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

Type

Clear all

Clear all

Fairness

Origin

Scope

SUBMIT A TOOL

If you have a tool that you think should be featured in the Catalogue of Tools & Metrics for Trustworthy AI, we would love to hear from you!

Submit
Objective Fairness
Approach Technical

TechnicalUploaded on Sep 27, 2026
EuConform is an open-source toolkit and open specification for producing technical evidence of compliance with the EU AI Act. It scans an AI system's code and produces machine-readable evidence files, including an AI bill of materials and a report of compliance gaps. Files are bundled with cryptographic hashes so they can be verified. It also offers AI Act risk classification and offline bias testing of language models. It provides technical guidance only, not legal advice.

TechnicalUploaded on Sep 27, 2026
PrivFair is a library for auditing the fairness of machine learning models while protecting privacy. Fairness audits require sensitive data about attributes such as gender or race, which is protected by law. PrivFair uses Secure Multiparty Computation, a cryptographic technique, so that an external auditor can test a company's model without either party revealing their data or model. It supports group fairness audits on tabular and image data. This enables independent, privacy-preserving checks of proprietary AI models.

Objective(s)

Related lifecycle stage(s)

Verify & validate

TechnicalNetherlandsUploaded on Sep 27, 2026
DP-CGANs is an open-source Python library for generating synthetic data that protects individual privacy, designed with personal health data in mind. It was developed at Maastricht University. The library uses a conditional generative adversarial network, adapted to capture relationships between variables in the data. Differential privacy can be enabled during training, so no single person strongly influences the result. It works with tabular and RDF data. It is still under development and released under the MIT licence.

TechnicalItalyUploaded on Sep 27, 2026
MANILA is a web-based, low-code application for benchmarking machine learning models and methods that reduce bias. It helps users find the combination that offers the best balance between fairness and performance. MANILA models a fairness benchmarking workflow as a set of options, such as datasets, models and metrics. Rules between these options guide users and prevent experiments that would fail. This makes fairness benchmarking accessible to users with little or no programming experience.

TechnicalUploaded on Sep 27, 2026
Fairmetrics is an open-source R package for evaluating the group fairness of machine learning models. It was developed by researchers at the University of Toronto. The package calculates a wide range of fairness metrics, such as statistical parity and equal opportunity. Unlike most fairness tools, it also provides confidence intervals for each metric. This shows whether differences between groups are statistically significant or due to random variation. It is available on CRAN under the MIT licence.

Related lifecycle stage(s)

Operate & monitorVerify & validate

TechnicalProceduralUploaded on Sep 27, 2026
IEEE 3198-2025 is a technical standard that specifies a method for evaluating the fairness of machine learning systems. It was published by the IEEE Standards Association in May 2025. The standard categorises the causes of unfairness in machine learning and presents widely used definitions of fairness. For each definition, it specifies metrics and explains how to calculate them. Test cases set out detailed procedures for fairness evaluations. This supports consistent, repeatable fairness testing across systems and organisations.

TechnicalSwedenUploaded on Sep 27, 2026
FairX is an open-source toolkit for benchmarking machine learning models on fairness, data utility and explainability. It was developed by researchers at Linköping University. Users load tabular or image datasets and apply bias-mitigation methods before, during or after model training, including fair generative models that create synthetic data. FairX evaluates results using fairness, performance and synthetic data quality metrics. This lets users compare trade-offs between fairness and utility. It is written in Python and released under the MIT licence.

TechnicalUploaded on Sep 27, 2026
Google's Differential Privacy libraries are an open-source collection of tools for producing statistics from sensitive data while protecting individuals' privacy. They add calibrated random noise to results such as counts and averages, so no single person can be identified. Building block libraries are available in C++, Go and Java. End-to-end frameworks make the tools usable by non-experts on large-scale data systems. Further tools track privacy budgets and audit privacy guarantees. The libraries are released under the Apache 2.0 licence.

TechnicalUploaded on Sep 26, 2026
Audit the Judge is an open-source tool for checking whether large language models used as judges evaluate answers reliably. It builds answer pairs where the correct verdict is known. These test whether the judge invents preferences between equal answers or favours answers shown first or written at greater length. Results are presented in a one-page report with statistical confidence intervals. The tool works with any major LLM judge. It is written in Python and released under the MIT licence.

TechnicalIndiaUploaded on Sep 26, 2026
ASR-FairBench is an online leaderboard that ranks speech recognition systems on both accuracy and fairness. It was developed by researchers at IIT Kharagpur. The leaderboard tests models on recordings from speakers of different ages, genders, ethnicities and backgrounds. It measures how error rates differ across these groups and combines this with overall accuracy into a single score. Models are rated from "severely biased" to "exemplarily fair". Users can submit their own models and receive a fairness audit within minutes.

TechnicalUnited StatesItalyUploaded on Sep 26, 2026
FairnessEval is an open-source framework for evaluating the fairness of machine learning models. It was developed by the University of Modena and Reggio Emilia with Microsoft. Users can load real or synthetic datasets and run fairness-aware models from toolkits such as Fairlearn and AIF360. The framework compares models on fairness, accuracy and training time, and presents the trade-offs in charts. It supports both selecting a suitable model and validating its performance under changing conditions.

TechnicalUnited StatesUploaded on Sep 26, 2026
CEB (Compositional Evaluation Benchmark for Fairness) is a fairness evaluation benchmark containing about 11 000 samples spanning multiple bias types, social groups, and tasks. It enables consistent measurement of compositional fairness behavior to identify failures that occur when fairness conditions combine across groups and bias factors. It is used by developers and auditors to support the fairness and accountability objectives, and to assess robustness of fairness performance under varied evaluation settings.

TechnicalKoreaUploaded on Sep 26, 2026
FLEX is a benchmark for testing whether large language models stay fair when prompts are designed to induce bias. It was developed by researchers at Korea University. The benchmark contains 3,145 multiple-choice questions from established fairness datasets, each combined with an adversarial prompt. These include assigning negative personas, forbidding refusals and introducing small text changes. Results show that models appearing fair on standard benchmarks can still be easily manipulated. The data and code are available for research purposes.

TechnicalUnited StatesUploaded on Sep 26, 2026
M4 Bias Eval FairFace is a dataset for evaluating social bias in vision-language models. It was created by Hugging Face to assess its IDEFICS models. It contains around 11,000 face images labelled by perceived gender, ethnicity and age. For each image, two models wrote a résumé, a dating profile and a news article about an arrest. Comparing these texts across demographic groups reveals stereotypes the models have learned. The dataset is released under a CC BY 4.0 licence.

TechnicalUploaded on Sep 25, 2026
IndoBias is a benchmark for evaluating social bias in large language models in Indonesian and three local languages. It was developed by researchers at the Mohamed bin Zayed University of Artificial Intelligence and Universitas Indonesia. One track tests whether models prefer stereotypical sentences over neutral ones. A second track tests whether models describe local groups and institutions more positively or negatively. The results show strong stereotypical bias in current models, especially regarding ideology and religion in local languages.

TechnicalUnited StatesUploaded on Sep 25, 2026
Microsoft's Responsible AI Toolbox is an open-source suite of tools for assessing and debugging machine learning models. It was developed by Microsoft. Its central dashboard combines error analysis, fairness assessment, model interpretability, counterfactual analysis and causal analysis in one interface. Users can identify where a model underperforms, understand why, and explore how outcomes could change. The toolbox supports models built on tabular data, text and images. It is released under the MIT licence.

TechnicalIndiaUploaded on Sep 25, 2026
FairMind is an open-source platform for AI governance and assurance, currently in development. It aims to help organisations evaluate AI systems for bias and collect reliable evidence for governance. Existing features test machine learning models against fairness metrics such as demographic parity. A new foundation records evaluation evidence with its scope, source and review status. Planned features include evaluation of large language models and mappings to frameworks such as the EU AI Act. The code is released under the MIT licence.

TechnicalKoreaUploaded on Sep 23, 2026
KSAFE-MM is a Korean-language benchmark for evaluating safety risks in multimodal large language models. It was developed by K-intelligence at KT in South Korea. The benchmark contains 14,135 query-image pairs across 11 risk categories, including hate, violence, privacy and weaponisation. One subset tests Korean culture-specific risks and applies jailbreak strategies such as role-play. Researchers use it to check whether models respond safely. It is available for research under a CC BY-NC 4.0 licence.

TechnicalFranceUploaded on Sep 22, 2026
DebiAI is an open-source web application that helps machine learning teams identify biases and errors in their data and model results. It supports data scientists in comparing model performance against contextual data, building custom dataset selections for further analysis or retraining, and creating shareable statistical visualisations through a customisable web dashboard.

TechnicalPolandUploaded on Sep 22, 2026
Fairmodels is an open-source R package for detecting, visualising and mitigating bias in machine learning classification and regression models. Using DALEX model explainers, it calculates twelve confusion-matrix-based fairness metrics across protected subgroups, summarised via a "parity loss" score, and offers pre- and post-processing techniques to reduce identified bias.

Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.