Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

Type

Clear all

Digital Security

Origin

Scope

SUBMIT A TOOL

If you have a tool that you think should be featured in the Catalogue of Tools & Metrics for Trustworthy AI, we would love to hear from you!

Submit
Objective Digital Security

TechnicalUploaded on Oct 7, 2026
Mandare is an open-source accountability system for fleets of AI agents. It gives each agent a signed identity and a signed mandate setting its spending caps, permissions and approval requirements. A gateway enforces these limits outside the agent, with an offline kill switch. Every action, including refusals, is recorded in a tamper-evident ledger that an independent witness can check. Raw data never leaves the user's machine. It is released under the AGPL-3.0 and Apache 2.0 licences.

TechnicalProceduralUploaded on Oct 2, 2026
Noisegate is an open-source demonstration of a gateway that lets AI agents query sensitive data without exposing individual records. Its privacy protections do not depend on the AI being trustworthy. A trusted layer checks every query against a policy, adds calibrated noise to answers and limits each user's privacy budget. Working attack demonstrations show its protections in action. The developer describes it as a demonstration, not a production product. It is released under the Apache 2.0 licence.

TechnicalIrelandUploaded on Oct 2, 2026
Diffprivlib is an open-source Python library for differential privacy, developed by IBM Research in Dublin. It adds calibrated random noise to data analysis and machine learning, so that no individual can be identified. The library provides privacy mechanisms, machine learning models with built-in privacy, data analysis tools and a privacy budget tracker. Its models work like those of scikit-learn, making them easy to adopt. It is intended for research and education, and was archived in September 2026.

TechnicalUploaded on Sep 23, 2026
Inner Warden is an open-source security agent for Linux and macOS servers. It detects attacks such as brute-force attempts and privilege escalation, and alerts operators in real time. It can use AI models to recommend responses, but AI remains advisory unless operators allow automatic action. Inner Warden can also monitor autonomous AI agents and block risky commands. All actions are reversible and recorded in an audit trail. It runs locally and is released under the MIT licence.

TechnicalUnited StatesUploaded on Sep 23, 2026
Sponsio is an open-source runtime safety tool for AI agents. It checks every agent action against deterministic rules, called agent contracts, before the action is executed. Contracts are based on formal methods and can allow, block, escalate or redirect actions. Every decision is logged in an audit trail. Users can apply ready-made contract bundles or draft rules in plain English. Sponsio works with major agent frameworks in Python and TypeScript, under the Apache 2.0 licence.

TechnicalUploaded on Sep 23, 2026
Doberman is an open-source security layer for AI coding agents such as Claude Code and Codex. It checks every action an agent takes before it runs. Routine actions pass, sensitive actions need human approval and dangerous actions are blocked. Any uncertainty results in the action being denied. Protections can tighten automatically, but weakening them requires human approval and is logged. Doberman is written in Python and released under the Apache 2.0 licence.

TechnicalUploaded on Sep 26, 2026
Gate AI: LLM Security Benchmark Evaluation Methodology and Results is a technical report on evaluating detectors of prompt injection and jailbreak attacks. It was published by Constellation Network. The method tests a detector across 16 public benchmarks with more than 12 000 samples. It uses a single detection threshold for all benchmarks and prevents near-duplicate prompts from inflating results. Competing detectors are compared at the same false-alarm rate. The report applies the method to Gate AI, Constellation Network's commercial security gateway.

TechnicalUploaded on Sep 27, 2026
Google's Differential Privacy libraries are an open-source collection of tools for producing statistics from sensitive data while protecting individuals' privacy. They add calibrated random noise to results such as counts and averages, so no single person can be identified. Building block libraries are available in C++, Go and Java. End-to-end frameworks make the tools usable by non-experts on large-scale data systems. Further tools track privacy budgets and audit privacy guarantees. The libraries are released under the Apache 2.0 licence.

TechnicalUnited StatesUploaded on Sep 27, 2026
The OpenDP Library is an open-source, modular collection of algorithms for differential privacy. It is the core library of the OpenDP Project, a community effort led from Harvard University. Users build analyses from small components that transform data or add calibrated noise to protect individuals. The library tracks the overall privacy cost of each analysis. It is written in Rust, with bindings for Python and R. It is still under development and released under the MIT licence.

TechnicalUnited StatesUploaded on Sep 27, 2026
The Safety Evaluation Benchmark for Vision LLMs is an open dataset for testing the safety and robustness of models that answer questions about images. It was developed by researchers led by the University of California, Santa Cruz. One part tests how models handle unusual images, such as sketches, and misleading questions. A second part tests resistance to attacks, including altered images and jailbreak prompts. The dataset contains about 2 000 samples and is released under the Apache 2.0 licence.

TechnicalUnited StatesUploaded on Sep 22, 2026
Meilynx is a software tool for governing large language models (LLMs) and AI agents used in regulated sectors such as financial services, insurance, healthcare, and human resources. It runs as a proxy between applications and model providers, where it redacts sensitive data, restricts model use, and controls agents' tool calls. Each decision is recorded in a tamper-evident audit log that can be mapped to frameworks such as the EU AI Act. The tool provides evidence for controls but does not certify compliance.

TechnicalUnited StatesUploaded on Sep 16, 2026
FLARE-AI is an open-source AI flaw and incident reporting platform that lets anyone document a flaw in any AI system and route a single, standardised report to multiple developers and coordinators at once. It enables any AI actor to document vulnerabilities, biases, or incidents and route a single, standardised (JSON-LD) report to multiple developers, coordinators, and registries in the ecosystem.

Related lifecycle stage(s)

Operate & monitorVerify & validate

ProceduralUploaded on Sep 14, 2026
The AIM Framework (Awareness, Identification, Mitigation) presents a stepwise approach for the implementation of risk management strategies. The framework is intended for AI developers working in private, academic, or public sectors. It features a checklist with indicative scenarios for awareness-raising and training purposes.

TechnicalUnited StatesUploaded on Sep 14, 2026
A conformance corpus and reference verifier for execution evidence about AI agents. Each vector is a signed attestation with an expected verdict, so an implementer can run someone else's verifier against the corpus and find out whether it accepts what it must accept and refuses what it must refuse. The corpus includes adversarial cases that a permissive verifier passes and a correct one rejects. It runs offline and requires no network access or account. Apache-2.0.

TechnicalUploaded on Sep 15, 2026
OpenART is an open-source framework designed to evaluate the security of AI agents in dynamic, long-horizon, and stateful environments. It stress-tests agent runtimes against multi-step state poisoning, privilege escalation, and tool-use vulnerabilities across more than 10,000 benchmark scenarios.

TechnicalAustraliaUploaded on Aug 21, 2026
GovAI is the official Australian Government technology AI platform. Announced in July 2025, GovAI offers Australian Public Service (APS) staff secure access to advanced AI capabilities through government-controlled infrastructure. GovAI provides technical services for Australian Government agencies building AI solutions. Our hosting environments and model access services help technical teams experiment, develop and deploy generative AI applications in a secure, compliant environment.

TechnicalUnited StatesUploaded on Sep 14, 2026
Corsair is an open source integration layer that lets AI agents securely connect to third-party apps (Gmail, Slack, Notion, GitHub, HubSpot, and more) without teams building custom OAuth and API handling. It enforces permission gating on sensitive actions, requiring human approval before an agent executes things like sending emails, and isolates credentials so agents never see raw API keys. Available hosted or fully self-hosted, giving developers accountability and control over how AI systems act on users' behalf.

TechnicalIndiaUploaded on Sep 14, 2026
Provael is an open source tool that red teams vision language action policies, the models that convert camera input and instructions into physical robot actions, by running adversarial attacks in simulation and reporting an attack success rate with a statistical confidence interval and a benign control. Results are issued as machine readable evidence, mapped to an independently authored embodied AI security taxonomy and cross referenced to current regulatory frameworks including the EU AI Act, the EU Machinery Regulation, and ISO 10218. The full testing functionality is available at no cost under the Apache 2.0 licence, and the project publishes, alongside its findings, an explicit account of which attack families have been validated against real models and which remain unvalidated.

TechnicalUnited StatesUploaded on Sep 18, 2026
Microsoft Agent Governance Toolkit is an open-source framework for governing AI agents at scale. The toolkit provides identity management, policy enforcement, authorization, and audit capabilities, helping organizations control agent actions and maintain accountability across agent workflows and enterprise systems.

TechnicalUploaded on Sep 21, 2026
Flakestorm is an open-source tool for testing the robustness and resilience of AI agents. It enables organisations to assess how autonomous systems respond to adversarial inputs, operational failures, and unexpected scenarios through automated stress testing, red teaming, and failure injection. The tool supports AI assurance efforts by helping identify vulnerabilities and improve the reliability and safety of AI agents prior to deployment.

Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.