Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

Type

Origin

Scope

SUBMIT A TOOL

If you have a tool that you think should be featured in the Catalogue of Tools & Metrics for Trustworthy AI, we would love to hear from you!

Submit

TechnicalUnited StatesUploaded on Sep 16, 2026
FLARE-AI is an open-source AI flaw and incident reporting platform that lets anyone document a flaw in any AI system and route a single, standardised report to multiple developers and coordinators at once. It enables any AI actor to document vulnerabilities, biases, or incidents and route a single, standardised (JSON-LD) report to multiple developers, coordinators, and registries in the ecosystem.

Related lifecycle stage(s)

Operate & monitorVerify & validate

TechnicalUploaded on Sep 14, 2026
Toolkit for evaluating fairness and bias in machine learning models using multiple subgroup fairness metrics (including parity and equalized-odds-style measures). It supports fairness auditing by quantifying disparities across demographic or other defined subgroups. Data scientists and developers can use it to verify and validate fairness properties and to guide improvements toward fairer model behavior.

Related lifecycle stage(s)

Operate & monitorVerify & validate

ProceduralEuropean UnionUploaded on Sep 14, 2026
The AIM Framework (Awareness, Identification, Mitigation) presents a stepwise approach for the implementation of risk management strategies. The framework is intended for AI developers working in private, academic, or public sectors. It features a checklist with indicative scenarios for awareness-raising and training purposes.

TechnicalUnited KingdomUploaded on Sep 8, 2026
CXO Ready is a commercial SaaS platform that helps organisations inventory their AI systems, score each one against the EU AI Act, UK GDPR, and ISO 42001, and generate a prioritised, evidence-backed action plan.

Related lifecycle stage(s)

Operate & monitorDeploy

TechnicalUploaded on Sep 3, 2026
explainX/explainx is a Python toolkit for generating explanations and debugging insights for black-box machine learning models. It produces explanation outputs for model predictions or behavior, suppo...

EducationalJapanUploaded on Sep 4, 2026
The AI Slop Side Effect Database documents indirect harms to legitimate users, creators, researchers, and organisations caused by the proliferation of low-quality AI-generated content and by countermeasures introduced to control it. It classifies cases across gatekeeping failures, content contamination, discriminatory bias, institutional invisibility, and service self-contamination, with evidence levels, affected parties, sources, and analytical commentary.

Related lifecycle stage(s)

Operate & monitor

TechnicalUploaded on Sep 17, 2026
SB-Bench is an open source benchmark that evaluates stereotype bias in Large Multimodal Models, using 7 500 real world images across nine social bias categories. It tests whether models default to stereotypical assumptions in ambiguous scenarios and finds that adding vision capabilities to language models consistently increases bias compared with text only baselines. The benchmark, code and dataset are publicly available, offering developers and auditors a reproducible tool for assessing fairness in multimodal AI systems.

Objective(s)

Related lifecycle stage(s)

Verify & validate

TechnicalUploaded on Sep 21, 2026
LangFair is a Python library for conducting use-case level LLM bias and fairness assessments

Objective(s)

Related lifecycle stage(s)

Verify & validateBuild & interpret model

TechnicalFranceUploaded on Sep 22, 2026
DebiAI is an open-source web application that helps machine learning teams identify biases and errors in their data and model results. It supports data scientists in comparing model performance against contextual data, building custom dataset selections for further analysis or retraining, and creating shareable statistical visualisations through a customisable web dashboard.

TechnicalPolandUploaded on Sep 22, 2026
Fairmodels is an open-source R package for detecting, visualising and mitigating bias in machine learning classification and regression models. Using DALEX model explainers, it calculates twelve confusion-matrix-based fairness metrics across protected subgroups, summarised via a "parity loss" score, and offers pre- and post-processing techniques to reduce identified bias.

ProceduralSpainUploaded on Sep 17, 2026
An open-source tool that audits the human-AI interaction layer of a decision-support AI system: how it presents its results, whether the person can correct it, and whether its alerts fire at the right moment. It scores that layer against Microsoft's HAX-18 and Google's PAIR design guidelines and returns concrete, evidence-anchored findings mapped to the EU AI Act and the NIST AI RMF.

TechnicalUploaded on Jun 3, 2026
AI Ethics for Fairness is a software application that supports the detection, evaluation and mitigation of bias in AI models by analysing datasets, training models and applying fairness processing techniques.

TechnicalUploaded on Jun 3, 2026
FACTS Grounding is a comprehensive benchmark for evaluating the ability of LLMs to generate responses that are not only factually accurate with respect to given inputs, but also sufficiently detailed to provide satisfactory answers to user queries.

Objective(s)

Related lifecycle stage(s)

Verify & validateBuild & interpret model

TechnicalProceduralUploaded on Jun 3, 2026
AuditNLG is an open-source toolkit for auditing the trustworthiness of generative AI text. It evaluates outputs across three key dimensions: factualness (consistency with knowledge), safety (harmful or biased content), and constraint adherence (compliance with instructions). The tool aggregates multiple state-of-the-art methods and provides scores, explanations, and improved text suggestions via self-refinement prompts. It supports both API-based and local models, enabling flexible integration into evaluation pipelines and governance frameworks.

ProceduralUploaded on Mar 20, 2026
Judgment Assurance is a decision-governance discipline that reframes human judgment as a governed institutional asset. It provides a structured framework and practical instruments, including the Underwriting Questionnaire (JA-UQ) and Maturity Model (JAMM-PS), to ensure that consequential AI-mediated decisions are reconstructible and defensible. By defining minimum governance controls for human oversight, it closes the "accountability gap," allowing institutions to define, record, own, and guard the reasoning behind consequential AI-supported outcomes.

TechnicalProceduralUploaded on Mar 20, 2026
The Approved Intelligence Platform (AIP) provides modular, scenario-based testing workflows to evaluate mission-critical AI systems in defence, public safety, and critical civil use cases. It delivers a comprehensive, end-to-end testing environment based on a proprietary AI trust ontology with measurable AI Solutions Quality Indicators (ASQI) for the testing, evaluation, validation and verification of software solutions with different AI modalities.

Objective(s)

Related lifecycle stage(s)

Operate & monitorDeployVerify & validate

ProceduralUploaded on Mar 20, 2026
The AI Governance Playbook from the Council on AI Governance helps organizations align people, processes and tools to achieve responsible AI outcomes.

ProceduralCanadaUploaded on Apr 1, 2025
An artificial intelligence (AI) impact assessment tool to provide organisations a method to assess AI systems for compliance with Canadian human rights law. The purpose of this human rights AI impact assessment is to assist developers and administrators of AI systems to identify, assess, minimise or avoid discrimination and uphold human rights obligations throughout the lifecycle of an AI system.

TechnicalUnited StatesUploaded on Mar 24, 2025
An open-source Python library designed for developers to calculate fairness metrics and assess bias in machine learning models. This library provides a comprehensive set of tools to ensure transparency, accountability, and ethical AI development.

TechnicalUnited StatesUploaded on Nov 8, 2024
The Python Risk Identification Tool for generative AI (PyRIT) is an open access automation framework to empower security professionals and machine learning engineers to proactively find risks in their generative AI systems.

Related lifecycle stage(s)

Operate & monitorVerify & validate

Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.