Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

Flakestorm



Flakestorm is an open-source testing and validation tool designed to assess the robustness and reliability of AI agents. It applies principles from chaos engineering and adversarial testing to evaluate how agents behave under challenging conditions, such as unexpected inputs, tool failures, latency issues, and security attacks. By systematically stress-testing AI systems, Flakestorm helps developers identify vulnerabilities and failure modes that may not be uncovered through conventional testing or benchmark evaluations. 

The tool works by generating adversarial mutations and injecting controlled disruptions into agent workflows, enabling organisations to observe how agents respond to uncertainty, errors, and malicious inputs. It supports activities such as agentic unit testing, red teaming, and resilience assessments, producing evidence on whether an agent behaves consistently and adheres to expected constraints under real-world conditions.

Auto-discovered on 2026-07-31 by OECD Catalogue Automation

Use Cases

There is no use cases for this tool yet.

Would you like to submit a use case for this tool?

If you have used this tool, we would love to know more about your experience.

Add use case
Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.