Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

Type

Clear all

Transparency

Origin

Scope

SUBMIT A TOOL

If you have a tool that you think should be featured in the Catalogue of Tools & Metrics for Trustworthy AI, we would love to hear from you!

Submit
Objective Transparency

ProceduralUnited KingdomUploaded on Oct 6, 2026
The CREST Accreditation Standards set requirements for organisations that provide cyber security services, verified through independent assessment. They are published by CREST, an international not-for-profit body. Three components address AI. Domain 7 requires all accredited organisations to govern their use of AI, including validating AI-supported outputs. Annex B sets requirements for using AI in penetration testing, with human validation of findings. A third standard covers security testing of generative AI and large language model-enabled systems.

TechnicalUploaded on Oct 2, 2026
DelusionEval is a dataset for evaluating how AI chatbots behave when conversations reinforce users' delusional beliefs. It was developed by Stanford University. It contains 725 excerpts from real, anonymised conversations, each labelled with one of 18 chatbot behaviours. These include harmful behaviours, such as endorsing delusions or facilitating self-harm, and protective ones, such as discouraging violence. Researchers use it to audit chatbot safety and develop detection methods. Access is restricted to non-commercial research.

EducationalProceduralUnited KingdomUploaded on Sep 30, 2026
AI Compliancy is an online EU AI Act risk assessment, obligation and report tool for UK small and medium-sized businesses that use AI built into everyday software. Users check whether the Act can reach them, record the software they use, assess what they use each AI feature for, work through the obligations that follow, and produce a dated report.

TechnicalUploaded on Sep 23, 2026
Inner Warden is an open-source security agent for Linux and macOS servers. It detects attacks such as brute-force attempts and privilege escalation, and alerts operators in real time. It can use AI models to recommend responses, but AI remains advisory unless operators allow automatic action. Inner Warden can also monitor autonomous AI agents and block risky commands. All actions are reversible and recorded in an audit trail. It runs locally and is released under the MIT licence.

TechnicalUnited StatesUploaded on Sep 26, 2026
M4 Bias Eval FairFace is a dataset for evaluating social bias in vision-language models. It was created by Hugging Face to assess its IDEFICS models. It contains around 11,000 face images labelled by perceived gender, ethnicity and age. For each image, two models wrote a résumé, a dating profile and a news article about an arrest. Comparing these texts across demographic groups reveals stereotypes the models have learned. The dataset is released under a CC BY 4.0 licence.

TechnicalUploaded on Sep 26, 2026
Gate AI: LLM Security Benchmark Evaluation Methodology and Results is a technical report on evaluating detectors of prompt injection and jailbreak attacks. It was published by Constellation Network. The method tests a detector across 16 public benchmarks with more than 12 000 samples. It uses a single detection threshold for all benchmarks and prevents near-duplicate prompts from inflating results. Competing detectors are compared at the same false-alarm rate. The report applies the method to Gate AI, Constellation Network's commercial security gateway.

TechnicalItalyUnited StatesUploaded on Sep 26, 2026
FairnessEval is an open-source framework for evaluating the fairness of machine learning models. It was developed by the University of Modena and Reggio Emilia with Microsoft. Users can load real or synthetic datasets and run fairness-aware models from toolkits such as Fairlearn and AIF360. The framework compares models on fairness, accuracy and training time, and presents the trade-offs in charts. It supports both selecting a suitable model and validating its performance under changing conditions.

TechnicalSwedenUploaded on Sep 27, 2026
FairX is an open-source toolkit for benchmarking machine learning models on fairness, data utility and explainability. It was developed by researchers at Linköping University. Users load tabular or image datasets and apply bias-mitigation methods before, during or after model training, including fair generative models that create synthetic data. FairX evaluates results using fairness, performance and synthetic data quality metrics. This lets users compare trade-offs between fairness and utility. It is written in Python and released under the MIT licence.

TechnicalUploaded on Sep 27, 2026
Fairmetrics is an open-source R package for evaluating the group fairness of machine learning models. It was developed by researchers at the University of Toronto. The package calculates a wide range of fairness metrics, such as statistical parity and equal opportunity. Unlike most fairness tools, it also provides confidence intervals for each metric. This shows whether differences between groups are statistically significant or due to random variation. It is available on CRAN under the MIT licence.

Related lifecycle stage(s)

Operate & monitorVerify & validate

TechnicalUnited StatesUploaded on Sep 22, 2026
Meilynx is a software tool for governing large language models (LLMs) and AI agents used in regulated sectors such as financial services, insurance, healthcare, and human resources. It runs as a proxy between applications and model providers, where it redacts sensitive data, restricts model use, and controls agents' tool calls. Each decision is recorded in a tamper-evident audit log that can be mapped to frameworks such as the EU AI Act. The tool provides evidence for controls but does not certify compliance.

ProceduralUnited StatesUploaded on Sep 22, 2026
AI Facts is a free, open-source Decision Terrain tool with an independently developed checklist mapped to NIST AI RMF 1.0. It is not affiliated with, sponsored by, or endorsed by NIST. Teams document self-reported safeguards, evidence notes, gaps, owners, and next actions, then export assessment reports and AI Facts labels. This public-alpha tool does not verify evidence or establish certification, compliance, or safety.

TechnicalUnited KingdomUploaded on Sep 22, 2026
The Model Card Builder is a free, browser based tool by HCXAIResearch for documenting AI systems that an organisation deploys but did not develop. Users complete eight sections covering provider documentation, intended use, stated capabilities and limitations, and their own oversight, safeguards and monitoring. Provider claims are kept separate from the organisation's own information. Cards can be tagged against the NIST AI Risk Management Framework and EU AI Act risk classes, previewed live, and exported as HTML, Markdown or PDF.

TechnicalSpainUploaded on Oct 7, 2026
contextburn is an open-source tool that measures how efficiently AI coding agents use the tokens they are paid for. Coding agents spend most tokens re-reading earlier context rather than producing new work. contextburn reads local Claude Code transcripts, without sending data over the network, and reports the share of tokens that became output. A second, cost-weighted measure reflects how users run their sessions.

Objective(s)

Related lifecycle stage(s)

Operate & monitor

TechnicalUploaded on Sep 14, 2026
Toolkit for evaluating fairness and bias in machine learning models using multiple subgroup fairness metrics (including parity and equalized-odds-style measures). It supports fairness auditing by quantifying disparities across demographic or other defined subgroups. Data scientists and developers can use it to verify and validate fairness properties and to guide improvements toward fairer model behavior.

Related lifecycle stage(s)

Operate & monitorVerify & validate

EducationalUploaded on Sep 14, 2026
The REFRAIME Legal Toolkit provides legal practitioners, public authorities, and civil society organisations with a structured resource for identifying and addressing the impact of AI systems on fundamental rights under the EU Charter. Developed by a consortium including the Center for the Study of Democracy, the European Center for Not-for-Profit Law, and the University of Malta, and co-funded by the European Union, the toolkit combines knowledge articles, sixteen real-world case studies grounded in actual case law (including ACLU v. Clearview AI, SCHUFA before the CJEU, and the Dutch Childcare Benefits case), an interactive glossary, a curated directory of EU, Council of Europe, UN, and OECD instruments, and a 43-point checklist for monitoring compliance with Fundamental Rights Impact Assessment obligations under Article 27 of the EU AI Act.

ProceduralUploaded on Sep 14, 2026
The AIM Framework (Awareness, Identification, Mitigation) presents a stepwise approach for the implementation of risk management strategies. The framework is intended for AI developers working in private, academic, or public sectors. It features a checklist with indicative scenarios for awareness-raising and training purposes.

TechnicalUploaded on Sep 15, 2026
ModelBench is a benchmarking tool for running safety-focused evaluations on AI models and producing detailed reports on performance against the benchmark suite. It is intended for developers and researchers who need to verify and compare safety/robustness behavior before deployment and during ongoing evaluation. The outputs support auditing and model validation for safety-related trustworthiness objectives.

TechnicalGermanyUploaded on Sep 7, 2026
Legalithm is a free, open-source developer toolkit that brings EU AI Act compliance directly into the software development workflow. It classifies an AI system's risk tier under Regulation (EU) 2024/1689 and generates a dated, auditable compliance record. It also supports Article 50(2) transparency duties by watermarking AI-generated content (via C2PA credentials and pixel watermarking) and verifying such marks. A GitHub Action integration fails continuous integration builds when the compliance record drifts from the codebase or from regulatory deadlines, helping engineering teams catch compliance gaps early.

TechnicalUnited StatesUploaded on Sep 16, 2026
TrustyAI Explainability Toolkit is a software toolkit for generating, transforming, and managing explanations of AI model behavior. It helps practitioners produce explanation artifacts that support the trustworthiness objectives of explainability and transparency, enabling users to inspect and communicate how model outputs are derived.

EducationalUnited StatesUploaded on Sep 9, 2026
Human Approval Gate is a free, platform-neutral educational kit that helps leaders, educators, operators, and small teams define what a qualified person must check before AI-assisted work can affect a real decision or action. It includes a practical guide, printable worksheet, facilitator notes, and ten synthetic test cases. Its CLEAR test holds the consequence, names the reviewer and evidence, preserves accept, revise, reject, and escalate outcomes, and records the decision and recovery path.

Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.