Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

Fairmetrics



Fairmetrics is an open-source R package for evaluating the group fairness of machine learning (ML) models. ML models are increasingly used in high-stakes areas such as criminal justice, healthcare, finance and employment, where biased outcomes for groups defined by attributes such as race, gender or age can reinforce inequalities. Existing fairness tools usually report single values for fairness metrics without showing how certain those values are. Users therefore cannot tell whether a difference between groups is statistically significant or simply due to random variation. fairmetrics addresses this gap by providing confidence intervals alongside each metric. It was developed by researchers at the University of Toronto.

The package covers binary classification tasks with two protected groups, such as loan approval or disease screening. It includes metrics from the three main families of group fairness. Independence metrics, such as statistical parity, compare how often each group receives a positive prediction. Separation metrics, such as equal opportunity and predictive equality, compare error rates across groups. Sufficiency metrics, such as predictive parity, compare how reliable positive and negative predictions are for each group. Further metrics compare overall accuracy and calibration.

Users provide a dataset containing the model's predicted probabilities, the true outcomes and the protected attribute. A single function then calculates all metrics, as both differences and ratios between groups. For each, it estimates confidence intervals using bootstrap resampling. An interval that excludes zero for a difference, or one for a ratio, indicates a statistically significant disparity. The package has no external dependencies, includes an example clinical dataset from the MIMIC-II database, and is available on CRAN under the MIT licence.

Auto-discovered on 2026-08-05 by OECD Catalogue Automation

Use Cases

There is no use cases for this tool yet.

Would you like to submit a use case for this tool?

If you have used this tool, we would love to know more about your experience.

Add use case
Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.