Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

Type

Clear all

Robustness

Origin

Scope

SUBMIT A TOOL

If you have a tool that you think should be featured in the Catalogue of Tools & Metrics for Trustworthy AI, we would love to hear from you!

Submit
Objective Robustness

TechnicalProceduralUnited KingdomUploaded on Sep 16, 2026
RecourseBench is a modular evaluation framework for algorithmic recourse methods that emphasises reproducibility when assessing user-facing counterfactual explanations. It enables practitioners to systematically compare recourse methods based on how they support actionable changes in decision-making systems. It is intended for researchers and developers who validate and improve the trustworthiness of explanation and human-agency features in AI used for consequential decisions.

Objective(s)

Related lifecycle stage(s)

Operate & monitorDeployVerify & validate

TechnicalUploaded on Sep 15, 2026
ModelBench is a benchmarking tool for running safety-focused evaluations on AI models and producing detailed reports on performance against the benchmark suite. It is intended for developers and researchers who need to verify and compare safety/robustness behavior before deployment and during ongoing evaluation. The outputs support auditing and model validation for safety-related trustworthiness objectives.

TechnicalUploaded on Sep 15, 2026
OpenART is an open-source framework designed to evaluate the security of AI agents in dynamic, long-horizon, and stateful environments. It stress-tests agent runtimes against multi-step state poisoning, privilege escalation, and tool-use vulnerabilities across more than 10,000 benchmark scenarios.

TechnicalUploaded on Sep 15, 2026
IndicSafeEval is a multilingual benchmark framework for evaluating the safety and robustness of large language models (LLMs) against persuasion-based jailbreak attacks in Indian languages. It combines safety-critical content categories with multiple human-like persuasion strategies and evaluates model responses across several languages. The framework is designed to identify safety and alignment failures in non-English settings and can be used to assess and compare model behavior, support model validation, and monitor safety performance.

TechnicalUploaded on Sep 16, 2026
Robust and Reliable Algorithmic Recourse (ROAR) is a framework for generating instance-level algorithmic recourse that is designed to remain reliable when the underlying predictive model changes. It helps practitioners evaluate and improve the robustness of recourse/decision-support explanations so that suggested actions continue to work under model or distribution shifts. Target users include researchers and developers working on fair/robust recourse systems to address robustness and accountability-related trustworthiness objectives.

TechnicalUnited StatesUploaded on Sep 16, 2026
Ster is an open-source framework for intervening in large language model behavior at the level of internal activations to reduce harmful outputs and hallucinations. It helps users improve safety and robustness by enabling targeted controls that affect what the model produces, supporting transparency by making activation-level intervention mechanisms available for inspection and explanation.

TechnicalUnited StatesUploaded on Sep 14, 2026
Corsair is an open source integration layer that lets AI agents securely connect to third-party apps (Gmail, Slack, Notion, GitHub, HubSpot, and more) without teams building custom OAuth and API handling. It enforces permission gating on sensitive actions, requiring human approval before an agent executes things like sending emails, and isolates credentials so agents never see raw API keys. Available hosted or fully self-hosted, giving developers accountability and control over how AI systems act on users' behalf.

TechnicalIndiaUploaded on Sep 14, 2026
Provael is an open source tool that red teams vision language action policies, the models that convert camera input and instructions into physical robot actions, by running adversarial attacks in simulation and reporting an attack success rate with a statistical confidence interval and a benign control. Results are issued as machine readable evidence, mapped to an independently authored embodied AI security taxonomy and cross referenced to current regulatory frameworks including the EU AI Act, the EU Machinery Regulation, and ISO 10218. The full testing functionality is available at no cost under the Apache 2.0 licence, and the project publishes, alongside its findings, an explicit account of which attack families have been validated against real models and which remain unvalidated.

TechnicalUploaded on Sep 21, 2026
Flakestorm is an open-source tool for testing the robustness and resilience of AI agents. It enables organisations to assess how autonomous systems respond to adversarial inputs, operational failures, and unexpected scenarios through automated stress testing, red teaming, and failure injection. The tool supports AI assurance efforts by helping identify vulnerabilities and improve the reliability and safety of AI agents prior to deployment.

TechnicalUnited StatesUploaded on Sep 21, 2026
NeMo Guardrails is an open-source toolkit for easily adding programmable guardrails to LLM-based conversational systems.

TechnicalEducationalProceduralUnited StatesUnited KingdomChinaInternationalEUUploaded on Aug 26, 2026<1 year
A three-layer, enterprise-wide AI risk management and governance framework that operationalises trustworthy AI from board strategy to operational controls and organisational resilience.

InternationalUploaded on Jun 4, 2026
Amazon Nova Premier is a multimodal foundation model that was evaluated under Amazon’s Frontier Model Safety Framework to assess and mitigate risks related to Chemical, Biological, Radiological, and Nuclear (CBRN) weapons proliferation, offensive cyber operations, and automated AI research and development.

TechnicalUploaded on Jun 3, 2026
ShieldGemma is a set of instruction tuned models for evaluating the safety of text and images against a set of defined safety policies.

Objective(s)

Related lifecycle stage(s)

Operate & monitorDeploy

TechnicalUploaded on Jun 3, 2026
FACTS Grounding is a comprehensive benchmark for evaluating the ability of LLMs to generate responses that are not only factually accurate with respect to given inputs, but also sufficiently detailed to provide satisfactory answers to user queries.

Objective(s)

Related lifecycle stage(s)

Verify & validateBuild & interpret model

TechnicalUploaded on Jun 3, 2026
Eureka is a reusable and open evaluation framework for standardizing evaluations of large foundation models beyond single-score reporting and rankings.

TechnicalUploaded on Jun 3, 2026
The MLCommons AILuminate benchmark evaluates an AI system-under-test (SUT) by inputting a set of prompts, recording the SUT’s responses, and then using a specialized set of “safety evaluators models” to determine which of the responses are violations according to the AILuminate Assessment Standard guidelines. Findings are summarized in a human-readable report.

Objective(s)

Related lifecycle stage(s)

Verify & validate

TechnicalSwedenUploaded on Sep 21, 2026
Pegasus is an open-source framework for automated compliance validation of AI systems against international standards. Uses 96 OPA Rego policies evaluated by a Rust-native engine to assess AI systems against ISO 42001, EU AI Act, NIST AI RMF, OWASP LLM Top 10, and 8 more standards. Features dual-agent cross-review with confidence scoring and content-addressable evidence store for audit trails.

TechnicalUploaded on Mar 20, 2026
garak is an open-source LLM vulnerability scanner developed by NVIDIA that probes large language models for security weaknesses including prompt injection, jailbreaks, hallucination, toxicity, data leakage, and misinformation.

TechnicalProceduralUploaded on Jun 3, 2026
AuditNLG is an open-source toolkit for auditing the trustworthiness of generative AI text. It evaluates outputs across three key dimensions: factualness (consistency with knowledge), safety (harmful or biased content), and constraint adherence (compliance with instructions). The tool aggregates multiple state-of-the-art methods and provides scores, explanations, and improved text suggestions via self-refinement prompts. It supports both API-based and local models, enabling flexible integration into evaluation pipelines and governance frameworks.

TechnicalUploaded on Mar 20, 2026
OpenEnv is a framework for evaluating AI agents against real systems rather than simulations. It provides a standardised way to connect agents to real tools and workflows while preserving the structure needed for consistent and reliable evaluation.

Objective(s)

Related lifecycle stage(s)

Operate & monitorDeployVerify & validate

Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.