Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

Type

Origin

Scope

SUBMIT A TOOL

If you have a tool that you think should be featured in the Catalogue of Tools & Metrics for Trustworthy AI, we would love to hear from you!

Submit

TechnicalUploaded on Sep 23, 2026
Inner Warden is an open-source security agent for Linux and macOS servers. It detects attacks such as brute-force attempts and privilege escalation, and alerts operators in real time. It can use AI models to recommend responses, but AI remains advisory unless operators allow automatic action. Inner Warden can also monitor autonomous AI agents and block risky commands. All actions are reversible and recorded in an audit trail. It runs locally and is released under the MIT licence.

TechnicalUnited StatesUploaded on Sep 23, 2026
Sponsio is an open-source runtime safety tool for AI agents. It checks every agent action against deterministic rules, called agent contracts, before the action is executed. Contracts are based on formal methods and can allow, block, escalate or redirect actions. Every decision is logged in an audit trail. Users can apply ready-made contract bundles or draft rules in plain English. Sponsio works with major agent frameworks in Python and TypeScript, under the Apache 2.0 licence.

TechnicalKoreaUploaded on Sep 23, 2026
KSAFE-MM is a Korean-language benchmark for evaluating safety risks in multimodal large language models. It was developed by K-intelligence at KT in South Korea. The benchmark contains 14,135 query-image pairs across 11 risk categories, including hate, violence, privacy and weaponisation. One subset tests Korean culture-specific risks and applies jailbreak strategies such as role-play. Researchers use it to check whether models respond safely. It is available for research under a CC BY-NC 4.0 licence.

TechnicalUploaded on Sep 23, 2026
Doberman is an open-source security layer for AI coding agents such as Claude Code and Codex. It checks every action an agent takes before it runs. Routine actions pass, sensitive actions need human approval and dangerous actions are blocked. Any uncertainty results in the action being denied. Protections can tighten automatically, but weakening them requires human approval and is logged. Doberman is written in Python and released under the Apache 2.0 licence.

TechnicalUploaded on Sep 25, 2026
IndoBias is a benchmark for evaluating social bias in large language models in Indonesian and three local languages. It was developed by researchers at the Mohamed bin Zayed University of Artificial Intelligence and Universitas Indonesia. One track tests whether models prefer stereotypical sentences over neutral ones. A second track tests whether models describe local groups and institutions more positively or negatively. The results show strong stereotypical bias in current models, especially regarding ideology and religion in local languages.

TechnicalUnited StatesUploaded on Sep 22, 2026
Meilynx is a software tool for governing large language models (LLMs) and AI agents used in regulated sectors such as financial services, insurance, healthcare, and human resources. It runs as a proxy between applications and model providers, where it redacts sensitive data, restricts model use, and controls agents' tool calls. Each decision is recorded in a tamper-evident audit log that can be mapped to frameworks such as the EU AI Act. The tool provides evidence for controls but does not certify compliance.

ProceduralUnited StatesUploaded on Sep 21, 2026
AI Facts is a free, open-source Decision Terrain tool with an independently developed checklist mapped to NIST AI RMF 1.0. It is not affiliated with, sponsored by, or endorsed by NIST. Teams document self-reported safeguards, evidence notes, gaps, owners, and next actions, then export assessment reports and AI Facts labels. This public-alpha tool does not verify evidence or establish certification, compliance, or safety.

TechnicalUnited StatesUploaded on Sep 16, 2026
FLARE-AI is an open-source AI flaw and incident reporting platform that lets anyone document a flaw in any AI system and route a single, standardised report to multiple developers and coordinators at once. It enables any AI actor to document vulnerabilities, biases, or incidents and route a single, standardised (JSON-LD) report to multiple developers, coordinators, and registries in the ecosystem.

Related lifecycle stage(s)

Operate & monitorVerify & validate

TechnicalProceduralUnited KingdomUploaded on Sep 16, 2026
RecourseBench is a modular evaluation framework for algorithmic recourse methods that emphasises reproducibility when assessing user-facing counterfactual explanations. It enables practitioners to systematically compare recourse methods based on how they support actionable changes in decision-making systems. It is intended for researchers and developers who validate and improve the trustworthiness of explanation and human-agency features in AI used for consequential decisions.

Objective(s)

Related lifecycle stage(s)

Operate & monitorDeployVerify & validate

TechnicalUnited KingdomUploaded on Sep 22, 2026
The Model Card Builder is a free, browser based tool by HCXAIResearch for documenting AI systems that an organisation deploys but did not develop. Users complete eight sections covering provider documentation, intended use, stated capabilities and limitations, and their own oversight, safeguards and monitoring. Provider claims are kept separate from the organisation's own information. Cards can be tagged against the NIST AI Risk Management Framework and EU AI Act risk classes, previewed live, and exported as HTML, Markdown or PDF.

TechnicalUploaded on Sep 14, 2026
Toolkit for evaluating fairness and bias in machine learning models using multiple subgroup fairness metrics (including parity and equalized-odds-style measures). It supports fairness auditing by quantifying disparities across demographic or other defined subgroups. Data scientists and developers can use it to verify and validate fairness properties and to guide improvements toward fairer model behavior.

Related lifecycle stage(s)

Operate & monitorVerify & validate

EducationalEuropean UnionUploaded on Sep 14, 2026
The REFRAIME Legal Toolkit provides legal practitioners, public authorities, and civil society organisations with a structured resource for identifying and addressing the impact of AI systems on fundamental rights under the EU Charter. Developed by a consortium including the Center for the Study of Democracy, the European Center for Not-for-Profit Law, and the University of Malta, and co-funded by the European Union, the toolkit combines knowledge articles, sixteen real-world case studies grounded in actual case law (including ACLU v. Clearview AI, SCHUFA before the CJEU, and the Dutch Childcare Benefits case), an interactive glossary, a curated directory of EU, Council of Europe, UN, and OECD instruments, and a 43-point checklist for monitoring compliance with Fundamental Rights Impact Assessment obligations under Article 27 of the EU AI Act.

ProceduralEuropean UnionUploaded on Sep 14, 2026
The AIM Framework (Awareness, Identification, Mitigation) presents a stepwise approach for the implementation of risk management strategies. The framework is intended for AI developers working in private, academic, or public sectors. It features a checklist with indicative scenarios for awareness-raising and training purposes.

TechnicalUnited StatesUploaded on Sep 14, 2026
A conformance corpus and reference verifier for execution evidence about AI agents. Each vector is a signed attestation with an expected verdict, so an implementer can run someone else's verifier against the corpus and find out whether it accepts what it must accept and refuses what it must refuse. The corpus includes adversarial cases that a permissive verifier passes and a correct one rejects. It runs offline and requires no network access or account. Apache-2.0.

TechnicalUploaded on Sep 15, 2026
SafeIMG is a safety-oriented benchmark for evaluating synthetic-image detection and visual-evidence verification in high-risk public- and personal-safety scenarios. It provides scenario-specific synthetic images generated for assessing how misleading visuals could impact safety and accountability. Researchers and developers can use it to verify and compare image authenticity methods, and to improve robustness of safeguards for trustworthy visual content.

Objective(s)

Related lifecycle stage(s)

Operate & monitorVerify & validate

TechnicalUploaded on Sep 15, 2026
ModelBench is a benchmarking tool for running safety-focused evaluations on AI models and producing detailed reports on performance against the benchmark suite. It is intended for developers and researchers who need to verify and compare safety/robustness behavior before deployment and during ongoing evaluation. The outputs support auditing and model validation for safety-related trustworthiness objectives.

TechnicalUploaded on Sep 15, 2026
OpenART is an open-source framework designed to evaluate the security of AI agents in dynamic, long-horizon, and stateful environments. It stress-tests agent runtimes against multi-step state poisoning, privilege escalation, and tool-use vulnerabilities across more than 10,000 benchmark scenarios.

TechnicalUploaded on Sep 15, 2026
IndicSafeEval is a multilingual benchmark framework for evaluating the safety and robustness of large language models (LLMs) against persuasion-based jailbreak attacks in Indian languages. It combines safety-critical content categories with multiple human-like persuasion strategies and evaluates model responses across several languages. The framework is designed to identify safety and alignment failures in non-English settings and can be used to assess and compare model behavior, support model validation, and monitor safety performance.

TechnicalGermanyUploaded on Sep 7, 2026
Legalithm is a free, open-source developer toolkit that brings EU AI Act compliance directly into the software development workflow. It classifies an AI system's risk tier under Regulation (EU) 2024/1689 and generates a dated, auditable compliance record. It also supports Article 50(2) transparency duties by watermarking AI-generated content (via C2PA credentials and pixel watermarking) and verifying such marks. A GitHub Action integration fails continuous integration builds when the compliance record drifts from the codebase or from regulatory deadlines, helping engineering teams catch compliance gaps early.

TechnicalUnited KingdomUploaded on Sep 8, 2026
CXO Ready is a commercial SaaS platform that helps organisations inventory their AI systems, score each one against the EU AI Act, UK GDPR, and ISO 42001, and generate a prioritised, evidence-backed action plan.

Related lifecycle stage(s)

Operate & monitorDeploy

Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.