These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.
Audit the Judge
Audit the Judge is an open-source tool for checking whether large language models (LLMs) used as judges evaluate answers reliably. Many AI benchmarks and leaderboards now use an LLM to decide which of two answers is better, a practice known as "LLM-as-a-judge". A biased judge can distort every ranking built on it, for example by inventing preferences between equally good answers or favouring whichever answer is shown first or is longer. The tool adapts a method originally developed for calibrating fairness audits, known as paired synthetic controls.
The tool builds pairs of answers where the correct verdict is already known. Negative controls are pairs of equivalent answers, where a reliable judge should not pick a winner. Positive controls test for known biases. Each pair is shown in both orders to detect position bias, and in concise and detailed versions to detect length bias. A further check confirms the judge can still distinguish a strong answer from a weak one.
Results are presented in a one-page Judge Trustworthiness Report. Every rate includes a statistical confidence interval, and the Benjamini–Hochberg false discovery rate correction ensures that biases flagged across many tests are not due to chance. To keep the test fair, the questions, the answers and the judgements are produced by separate models. The developer reports that all five major LLM judges tested preferred longer answers, to varying degrees. The tool works with any judge accessible through an OpenAI-compatible interface. It is written in Python and released under the MIT licence.
Auto-discovered on 2026-09-23 by OECD Catalogue Automation
About the tool
You can click on the links to see the associated tools
Tool type(s):
Objective(s):
Purpose(s):
Lifecycle stage(s):
Type of approach:
Maturity:
Usage rights:
License:
Target users:
Risk management stage(s):
Technology platforms:
Programming languages:
Github stars:
- 141
Github forks:
- 2
Use Cases
Would you like to submit a use case for this tool?
If you have used this tool, we would love to know more about your experience.
Add use case




























