Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

Type

Clear all

Data Governance & Traceability

Origin

Scope

SUBMIT A TOOL

If you have a tool that you think should be featured in the Catalogue of Tools & Metrics for Trustworthy AI, we would love to hear from you!

Submit
Objective Data Governance & Traceability

TechnicalLithuaniaUploaded on Oct 7, 2026
Agent Barn is an open-source control plane for running and governing AI agents on an organisation's own infrastructure. It was built by AAI Labs in Lithuania. Agents connect to team chat tools such as Slack and Microsoft Teams, and act on systems such as GitHub and Jira. Access roles control who manages each agent. Audit logs record conversations and tool calls, and costs are attributed to each agent. It is released under the Apache 2.0 licence.

TechnicalUploaded on Oct 7, 2026
Mandare is an open-source accountability system for fleets of AI agents. It gives each agent a signed identity and a signed mandate setting its spending caps, permissions and approval requirements. A gateway enforces these limits outside the agent, with an offline kill switch. Every action, including refusals, is recorded in a tamper-evident ledger that an independent witness can check. Raw data never leaves the user's machine. It is released under the AGPL-3.0 and Apache 2.0 licences.

TechnicalJapanUploaded on Oct 7, 2026
The VeritasChain Protocol (VCP) is an open specification for cryptographically verifiable audit trails in algorithmic and AI-driven trading. It is maintained by the VeritasChain Standards Organization. VCP lets regulators and auditors verify mathematically that records of trading decisions, orders, risk controls and AI model use are complete and unaltered. A privacy module allows personal data to be erased while keeping the audit trail intact. VCP is designed to support MiFID II, EU AI Act and GDPR compliance.

TechnicalIrelandUploaded on Oct 2, 2026
Diffprivlib is an open-source Python library for differential privacy, developed by IBM Research in Dublin. It adds calibrated random noise to data analysis and machine learning, so that no individual can be identified. The library provides privacy mechanisms, machine learning models with built-in privacy, data analysis tools and a privacy budget tracker. Its models work like those of scikit-learn, making them easy to adopt. It is intended for research and education, and was archived in September 2026.

EducationalProceduralUnited KingdomUploaded on Sep 30, 2026
AI Compliancy is an online EU AI Act risk assessment, obligation and report tool for UK small and medium-sized businesses that use AI built into everyday software. Users check whether the Act can reach them, record the software they use, assess what they use each AI feature for, work through the obligations that follow, and produce a dated report.

TechnicalUploaded on Sep 25, 2026
IndoBias is a benchmark for evaluating social bias in large language models in Indonesian and three local languages. It was developed by researchers at the Mohamed bin Zayed University of Artificial Intelligence and Universitas Indonesia. One track tests whether models prefer stereotypical sentences over neutral ones. A second track tests whether models describe local groups and institutions more positively or negatively. The results show strong stereotypical bias in current models, especially regarding ideology and religion in local languages.

TechnicalUploaded on Sep 26, 2026
Audit the Judge is an open-source tool for checking whether large language models used as judges evaluate answers reliably. It builds answer pairs where the correct verdict is known. These test whether the judge invents preferences between equal answers or favours answers shown first or written at greater length. Results are presented in a one-page report with statistical confidence intervals. The tool works with any major LLM judge. It is written in Python and released under the MIT licence.

TechnicalUnited StatesUploaded on Sep 27, 2026
The OpenDP Library is an open-source, modular collection of algorithms for differential privacy. It is the core library of the OpenDP Project, a community effort led from Harvard University. Users build analyses from small components that transform data or add calibrated noise to protect individuals. The library tracks the overall privacy cost of each analysis. It is written in Rust, with bindings for Python and R. It is still under development and released under the MIT licence.

TechnicalUnited StatesUploaded on Sep 27, 2026
dpmm is an open-source Python library for generating synthetic tabular data with differential privacy guarantees. It was developed by SAS. Synthetic data preserves the statistical patterns of real data without reproducing real individuals, so it can be shared or used to develop AI models. The library includes three widely used models: PrivBayes, MST and AIM. It protects privacy across the whole process, including data preparation, and addresses known vulnerabilities in differential privacy software.

TechnicalNetherlandsUploaded on Sep 27, 2026
DP-CGANs is an open-source Python library for generating synthetic data that protects individual privacy, designed with personal health data in mind. It was developed at Maastricht University. The library uses a conditional generative adversarial network, adapted to capture relationships between variables in the data. Differential privacy can be enabled during training, so no single person strongly influences the result. It works with tabular and RDF data. It is still under development and released under the MIT licence.

TechnicalUnited StatesUploaded on Sep 27, 2026
The SmartNoise SDK is an open-source toolkit for applying differential privacy to tabular and relational data. It was originally developed by Microsoft with Harvard's OpenDP initiative. SmartNoise SQL runs standard SQL queries and adds calibrated noise to results, so no individual can be identified. SmartNoise Synth generates synthetic data that preserves statistical patterns without reproducing real people. This lets organisations share data and develop AI models while protecting privacy. It is written in Python under the MIT licence.

TechnicalUnited StatesUploaded on Sep 22, 2026
Meilynx is a software tool for governing large language models (LLMs) and AI agents used in regulated sectors such as financial services, insurance, healthcare, and human resources. It runs as a proxy between applications and model providers, where it redacts sensitive data, restricts model use, and controls agents' tool calls. Each decision is recorded in a tamper-evident audit log that can be mapped to frameworks such as the EU AI Act. The tool provides evidence for controls but does not certify compliance.

ProceduralUnited StatesUploaded on Sep 22, 2026
AI Facts is a free, open-source Decision Terrain tool with an independently developed checklist mapped to NIST AI RMF 1.0. It is not affiliated with, sponsored by, or endorsed by NIST. Teams document self-reported safeguards, evidence notes, gaps, owners, and next actions, then export assessment reports and AI Facts labels. This public-alpha tool does not verify evidence or establish certification, compliance, or safety.

TechnicalUnited KingdomUploaded on Sep 22, 2026
The Model Card Builder is a free, browser based tool by HCXAIResearch for documenting AI systems that an organisation deploys but did not develop. Users complete eight sections covering provider documentation, intended use, stated capabilities and limitations, and their own oversight, safeguards and monitoring. Provider claims are kept separate from the organisation's own information. Cards can be tagged against the NIST AI Risk Management Framework and EU AI Act risk classes, previewed live, and exported as HTML, Markdown or PDF.

TechnicalUnited StatesUploaded on Sep 14, 2026
A conformance corpus and reference verifier for execution evidence about AI agents. Each vector is a signed attestation with an expected verdict, so an implementer can run someone else's verifier against the corpus and find out whether it accepts what it must accept and refuses what it must refuse. The corpus includes adversarial cases that a permissive verifier passes and a correct one rejects. It runs offline and requires no network access or account. Apache-2.0.

TechnicalGermanyUploaded on Sep 7, 2026
Legalithm is a free, open-source developer toolkit that brings EU AI Act compliance directly into the software development workflow. It classifies an AI system's risk tier under Regulation (EU) 2024/1689 and generates a dated, auditable compliance record. It also supports Article 50(2) transparency duties by watermarking AI-generated content (via C2PA credentials and pixel watermarking) and verifying such marks. A GitHub Action integration fails continuous integration builds when the compliance record drifts from the codebase or from regulatory deadlines, helping engineering teams catch compliance gaps early.

TechnicalUnited KingdomUploaded on Sep 8, 2026
CXO Ready is a commercial SaaS platform that helps organisations inventory their AI systems, score each one against the EU AI Act, UK GDPR, and ISO 42001, and generate a prioritised, evidence-backed action plan.

Related lifecycle stage(s)

Operate & monitorDeploy

EducationalJapanUploaded on Sep 4, 2026
The AI Slop Side Effect Database documents indirect harms to legitimate users, creators, researchers, and organisations caused by the proliferation of low-quality AI-generated content and by countermeasures introduced to control it. It classifies cases across gatekeeping failures, content contamination, discriminatory bias, institutional invisibility, and service self-contamination, with evidence levels, affected parties, sources, and analytical commentary.

Related lifecycle stage(s)

Operate & monitor

TechnicalUnited KingdomUploaded on Sep 4, 2026
The Green Algorithms calculator is an open-access online tool designed to estimate and report the carbon footprint of computational tasks and AI models. Developed by researchers at the University of Cambridge, the initiative addresses the growing, yet often overlooked, environmental impact of modern computing, ranging from high-performance scientific simulations to AI models and big data analytics. It can be used during the planning phase to estimate environmental impacts, or retrospectively for accounting and monitoring.

TechnicalEstoniaUploaded on Sep 14, 2026
An open specification and command line interface (CLI) toolkit for cryptographically pre-registering machine learning evaluation criteria, including the metric, threshold, dataset, and seed, before a model is run. By committing these criteria in advance, the tool makes any post hoc changes to success thresholds detectable rather than silent, supporting greater integrity and accountability in reported ML evaluation claims.

Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.