Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

OpenART



OpenART is an open-source framework for evaluating the security of autonomous AI agents in dynamic, long-horizon environments. It runs large-scale red-team stress tests that include multi-step poisoning, privilege escalation, and vulnerability scenarios. It is intended for developers and researchers to verify agent trustworthiness and to identify failure modes to improve agent design and runtime defenses.  

Within Docker environments, it runs AI agents with tasks that require multiple steps (e.g. knowledge-base integration, release syncing). It then introduces an "attacker" agent or adversarial configuration designed to manipulate the target agent mid-task, testing for vulnerabilities such as: 

  • State poisoning (corrupting the agent's memory/context over a long task)
  • Privilege escalation (agent being tricked into exceeding its intended permissions)
  • Tool-use exploitation (misuse or hijacking of the tools the agent calls) 

Evaluation in OpenART is conducted using deterministic and/or attacker-based strategies, with each run producing structured outputs that can be reviewed or scored afterward. Rather than relying solely on a fixed set of prebuilt scenarios, the framework includes a "planner" module that can generate new test tasks from a written scenario description, enabling evaluators to extend coverage beyond the bundled examples. Out of the box, it includes a small set of high-complexity example tasks involving up to approximately 100 workflow steps, more than 100 workspace files, and around 30 distinct tools each. These tasks are sufficient to demonstrate and run the framework without additional setup.

Auto-discovered on 2026-09-09 by OECD Catalogue Automation

Use Cases

There is no use cases for this tool yet.

Would you like to submit a use case for this tool?

If you have used this tool, we would love to know more about your experience.

Add use case
Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.