Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

Type

Clear all

Explainability

Origin

Scope

SUBMIT A TOOL

If you have a tool that you think should be featured in the Catalogue of Tools & Metrics for Trustworthy AI, we would love to hear from you!

Submit
Objective Explainability

TechnicalLithuaniaUploaded on Oct 7, 2026
Agent Barn is an open-source control plane for running and governing AI agents on an organisation's own infrastructure. It was built by AAI Labs in Lithuania. Agents connect to team chat tools such as Slack and Microsoft Teams, and act on systems such as GitHub and Jira. Access roles control who manages each agent. Audit logs record conversations and tool calls, and costs are attributed to each agent. It is released under the Apache 2.0 licence.

TechnicalIndiaUploaded on Oct 7, 2026
AiEnsured is a commercial testing platform for AI products, designed to support the deployment of responsible AI. It is developed in India. The platform brings together fairness and bias evaluation, explainability, adversarial robustness testing and privacy and GDPR compliance checks. It also generates test cases automatically, including rare and extreme inputs. Further features cover experiment management, model comparison and performance testing. After deployment, it monitors models for concept drift, which can reduce accuracy over time.

TechnicalSwedenUploaded on Sep 27, 2026
FairX is an open-source toolkit for benchmarking machine learning models on fairness, data utility and explainability. It was developed by researchers at Linköping University. Users load tabular or image datasets and apply bias-mitigation methods before, during or after model training, including fair generative models that create synthetic data. FairX evaluates results using fairness, performance and synthetic data quality metrics. This lets users compare trade-offs between fairness and utility. It is written in Python and released under the MIT licence.

TechnicalUnited KingdomUploaded on Sep 22, 2026
The Model Card Builder is a free, browser based tool by HCXAIResearch for documenting AI systems that an organisation deploys but did not develop. Users complete eight sections covering provider documentation, intended use, stated capabilities and limitations, and their own oversight, safeguards and monitoring. Provider claims are kept separate from the organisation's own information. Cards can be tagged against the NIST AI Risk Management Framework and EU AI Act risk classes, previewed live, and exported as HTML, Markdown or PDF.

TechnicalUploaded on Sep 15, 2026
IndicSafeEval is a multilingual benchmark framework for evaluating the safety and robustness of large language models (LLMs) against persuasion-based jailbreak attacks in Indian languages. It combines safety-critical content categories with multiple human-like persuasion strategies and evaluates model responses across several languages. The framework is designed to identify safety and alignment failures in non-English settings and can be used to assess and compare model behavior, support model validation, and monitor safety performance.

TechnicalUploaded on Sep 3, 2026
explainX/explainx is a Python toolkit for generating explanations and debugging insights for black-box machine learning models. It produces explanation outputs for model predictions or behavior, suppo...

TechnicalUploaded on Sep 16, 2026
Robust and Reliable Algorithmic Recourse (ROAR) is a framework for generating instance-level algorithmic recourse that is designed to remain reliable when the underlying predictive model changes. It helps practitioners evaluate and improve the robustness of recourse/decision-support explanations so that suggested actions continue to work under model or distribution shifts. Target users include researchers and developers working on fair/robust recourse systems to address robustness and accountability-related trustworthiness objectives.

TechnicalUnited StatesUploaded on Sep 16, 2026
TrustyAI Explainability Toolkit is a software toolkit for generating, transforming, and managing explanations of AI model behavior. It helps practitioners produce explanation artifacts that support the trustworthiness objectives of explainability and transparency, enabling users to inspect and communicate how model outputs are derived.

TechnicalUnited StatesUploaded on Sep 14, 2026
Open-source governance evidence workflow for machine-learning-based small-business credit underwriting. It provides schemas, templates, synthetic fixtures, public-data run kits, validation scripts, monitoring reports, adverse-action reason QA checks, model-change review, issue registers, and reviewer-ready evidence packs for model-risk, transparency, and accountable oversight review.

TechnicalUnited StatesUploaded on Sep 25, 2026
Microsoft's Responsible AI Toolbox is an open-source suite of tools for assessing and debugging machine learning models. It was developed by Microsoft. Its central dashboard combines error analysis, fairness assessment, model interpretability, counterfactual analysis and causal analysis in one interface. Users can identify where a model underperforms, understand why, and explore how outcomes could change. The toolbox supports models built on tabular data, text and images. It is released under the MIT licence.

ProceduralSpainUploaded on Sep 17, 2026
An open-source tool that audits the human-AI interaction layer of a decision-support AI system: how it presents its results, whether the person can correct it, and whether its alerts fire at the right moment. It scores that layer against Microsoft's HAX-18 and Google's PAIR design guidelines and returns concrete, evidence-anchored findings mapped to the EU AI Act and the NIST AI RMF.

ProceduralUploaded on Sep 23, 2026
AI Law Radar is an online tracker of AI laws and obligations across 79 jurisdictions. It consolidates them into a single register, with every entry linked to its primary source. Users select where they operate and what they do with AI. The tracker then shows which obligations apply to them and when. It includes a deadline calendar, a daily changelog and email alerts. The data is openly available under a CC BY 4.0 licence. It is not legal advice.

Related lifecycle stage(s)

Plan & design

TechnicalProceduralColombiaUploaded on Jun 9, 2026
Web application that allows organizations to assess their level of maturity in artificial intelligence governance and automatically generate a customized roadmap to meet national and international standards

TechnicalUploaded on Jun 3, 2026
SegMate is an open source AI Toolkit developed by the Vector Institute, which can help organizations and researchers apply cutting-edge computer vision techniques in the fight against climate change

TechnicalProceduralUploaded on Jun 3, 2026
AuditNLG is an open-source toolkit for auditing the trustworthiness of generative AI text. It evaluates outputs across three key dimensions: factualness (consistency with knowledge), safety (harmful or biased content), and constraint adherence (compliance with instructions). The tool aggregates multiple state-of-the-art methods and provides scores, explanations, and improved text suggestions via self-refinement prompts. It supports both API-based and local models, enabling flexible integration into evaluation pipelines and governance frameworks.

TechnicalUploaded on Jan 19, 2026
ASQI Engineer is an open-source framework for testing and assuring AI systems. Built for scale and reliability, it uses containerised test packages, automated assessments, and repeatable workflows to make evaluation transparent and robust. With ASQI Engineer, organisations also run ASQIs that they have created themselves, giving teams full control and confidence in AI quality.

TechnicalUploaded on Oct 9, 2025
An open-source framework for large language model evaluations. Inspect can be used for a broad range of evaluations that measure coding, agentic tasks, reasoning, knowledge, behavior, and multi-modal understanding.

Related lifecycle stage(s)

Operate & monitorVerify & validate

EducationalMaltaUploaded on Sep 1, 2025<1 day
This is a complete workshop package for the teaching of practical governance tools when using AI in collaborative teams. It covers the theoretical background, the facilitator notes for each phase, and the student workbook.

Uploaded on Aug 4, 2025
KOBI is a groundbreaking reading app specifically designed to support children with dyslexia. Grounded in robust scientific principles and a deep understanding of reading difficulties, KOBI integrates evidence-based methodologies to create an effective and engaging learning experience. Here is the science behind KOBI, highlighting its key features and the research that informs its development.

TechnicalUnited StatesUploaded on May 15, 2025
The GDA leverages aerial imagery, satellite data, and machine learning techniques to evaluate the damage in areas impacted by natural disasters. This tool greatly enhances the efficiency and precision of disaster response operations.

Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.