Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

kryptologyst/Fairness-Evaluation-Toolkit



The Fairness Evaluation Toolkit is an open-source Python library (MIT licensed) for assessing fairness and bias in machine learning classification models. It implements a set of standard fairness metrics (e.g. demographic parity difference, equalized odds difference, disparate impact ratio, statistical parity difference, equal opportunity difference, and calibration difference), enabling developers and researchers to quantify how a model's predictions vary across sensitive/protected subgroups.

The toolkit includes baseline model implementations (logistic regression, random forest, decision tree), synthetic dataset generators for creating controlled biased and unbiased test scenarios, and an interactive Streamlit dashboard for exploring fairness results in real time. It supports reproducible evaluation through deterministic seeding and provides a programmatic API alongside command-line scripts for running evaluation suites.

Notably, the toolkit incorporates explicit accountability and human-oversight safeguards. It includes a non-reliance disclaimer stating results should not be the sole basis for fairness decisions in production systems, and calls for continuous monitoring, multi-metric validation, documentation of evaluation processes, and regular fairness audits. It acknowledges known limitations, including metric instability, context-dependency of fairness definitions, and sensitivity to model complexity. 

Auto-discovered on 2026-08-13 by OECD Catalogue Automation

Use Cases

There is no use cases for this tool yet.

Would you like to submit a use case for this tool?

If you have used this tool, we would love to know more about your experience.

Add use case
Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.