Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

Audit the Judge



Audit the Judge is an open-source tool for checking whether large language models (LLMs) used as judges evaluate answers reliably. Many AI benchmarks and leaderboards now use an LLM to decide which of two answers is better, a practice known as "LLM-as-a-judge". A biased judge can distort every ranking built on it, for example by inventing preferences between equally good answers or favouring whichever answer is shown first or is longer. The tool adapts a method originally developed for calibrating fairness audits, known as paired synthetic controls.

The tool builds pairs of answers where the correct verdict is already known. Negative controls are pairs of equivalent answers, where a reliable judge should not pick a winner. Positive controls test for known biases. Each pair is shown in both orders to detect position bias, and in concise and detailed versions to detect length bias. A further check confirms the judge can still distinguish a strong answer from a weak one.

Results are presented in a one-page Judge Trustworthiness Report. Every rate includes a statistical confidence interval, and the Benjamini–Hochberg false discovery rate correction ensures that biases flagged across many tests are not due to chance. To keep the test fair, the questions, the answers and the judgements are produced by separate models. The developer reports that all five major LLM judges tested preferred longer answers, to varying degrees. The tool works with any judge accessible through an OpenAI-compatible interface. It is written in Python and released under the MIT licence.

Auto-discovered on 2026-09-23 by OECD Catalogue Automation

Use Cases

There is no use cases for this tool yet.

Would you like to submit a use case for this tool?

If you have used this tool, we would love to know more about your experience.

Add use case
Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.