Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

PRML : Pre-Registered ML Manifest Specification



PRML is an open specification and command-line toolkit for pre-registering machine learning evaluation claims before a model is run. A manifest records the metric, comparator, threshold, dataset identifier and seed for a claim, which is then canonicalised and hashed with SHA-256 to produce a fixed, tamper-evident reference. After the evaluation runs, the observed result is checked against the locked manifest: the tool returns a pass, fail, or "tampered" result if the criteria were altered after locking. The approach adapts the pre-registration discipline used in clinical trials to machine learning, aiming to prevent retroactive adjustment of success thresholds or silent re-running of evaluations until a desired result is obtained.

The specification is implemented in four reference languages (Python, JavaScript, Go, Rust) verified against a shared conformance suite, and integrates with common ML tooling such as MLflow, GitHub Actions, and several evaluation harnesses. It has been mapped by its author to provisions of the EU AI Act, the NIST AI Risk Management Framework, and ISO/IEC 42001.

The specification and reference implementations are released under open licences (Community Specification License 1.0 and MIT respectively). The instrument does not verify that a pre-registered evaluation was in fact the one executed, and independent (non-author) implementations or adoption are not yet established.

About the tool


Developing organisation(s):







Country/Territory of origin:



Type of approach:







Stakeholder group:





Geographical scope:





Technology platforms:



Tags:

  • ai assurance
  • pre-registration
  • sha-256
  • evaluation integrity
  • rfc 3161
  • conformance vectors

Github stars:

  • 7

Github forks:

  • 3

Modify this tool

Use Cases

There is no use cases for this tool yet.

Would you like to submit a use case for this tool?

If you have used this tool, we would love to know more about your experience.

Add use case
Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.