These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.
OpenART
OpenART is an open-source framework for evaluating the security of autonomous AI agents in dynamic, long-horizon environments. It runs large-scale red-team stress tests that include multi-step poisoning, privilege escalation, and vulnerability scenarios. It is intended for developers and researchers to verify agent trustworthiness and to identify failure modes to improve agent design and runtime defenses.
Within Docker environments, it runs AI agents with tasks that require multiple steps (e.g. knowledge-base integration, release syncing). It then introduces an "attacker" agent or adversarial configuration designed to manipulate the target agent mid-task, testing for vulnerabilities such as:
- State poisoning (corrupting the agent's memory/context over a long task)
- Privilege escalation (agent being tricked into exceeding its intended permissions)
- Tool-use exploitation (misuse or hijacking of the tools the agent calls)
Evaluation in OpenART is conducted using deterministic and/or attacker-based strategies, with each run producing structured outputs that can be reviewed or scored afterward. Rather than relying solely on a fixed set of prebuilt scenarios, the framework includes a "planner" module that can generate new test tasks from a written scenario description, enabling evaluators to extend coverage beyond the bundled examples. Out of the box, it includes a small set of high-complexity example tasks involving up to approximately 100 workflow steps, more than 100 workspace files, and around 30 distinct tools each. These tasks are sufficient to demonstrate and run the framework without additional setup.
Auto-discovered on 2026-09-09 by OECD Catalogue Automation
About the tool
You can click on the links to see the associated tools
Tool type(s):
Objective(s):
Impacted stakeholders:
Purpose(s):
Lifecycle stage(s):
Type of approach:
Maturity:
Usage rights:
License:
Target groups:
Target users:
Stakeholder group:
People involved:
Risk management stage(s):
Technology platforms:
Programming languages:
Tags:
- Application Security
- Identity & Access Management
- Vulnerability Assessment & Penetration Testing
Github stars:
- 171
Github forks:
- 13
Use Cases
Would you like to submit a use case for this tool?
If you have used this tool, we would love to know more about your experience.
Add use case




























