Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

Agent Evidence Conformance Suite



Agent Evidence Vectors is a technical conformance testing toolkit that provides a reference implementation and machine checkable test suite for verifying cryptographically signed evidence of AI agent execution. The tool implements two proposed in toto attestation predicates, Adversarial Execution Evidence and AI Agent Action, which specify how an AI agent's actions and their outcomes can be recorded and independently recomputed by a verifier, rather than accepted on the basis of a self-reported result.

The toolkit comprises over 270 conformance test vectors, a four-stage verification pipeline (statement well formedness, coverage validity, result recomputation, and evidence tier derivation), a standardised set of machine-readable failure codes, and a companion specification for binding evaluation records to auditable artifacts. It further supports transport of attestations via IETF SCITT/COSE transparency mechanisms and enables independent verification of the toolkit's own releases through timestamping and public blockchain anchored proofs.

The tool is intended for developers, system integrators and technical assessors who need to establish verifiable, auditable evidence of AI agent behaviour in production or evaluation settings. It supports conformance testing against a specification rather than a single implementation, allowing multiple independently developed verifiers to certify against the same corpus of test vectors. Cross implementation validation, including an independently developed verifier, is documented and published alongside the toolkit.

In the context of AI risk management, the tool addresses the assessment and governance stages. It provides a mechanism for producing and validating tamper evident execution records, supporting accountability, traceability, and auditability of AI agent actions, rather than defining policy, treating identified risks, or delivering an end user governance framework.

Use Cases

There is no use cases for this tool yet.

Would you like to submit a use case for this tool?

If you have used this tool, we would love to know more about your experience.

Add use case
Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.