Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

M4 Bias Eval FairFace



M4 Bias Eval FairFace is a dataset for evaluating social bias in vision-language models, which generate text from images. It was created by Hugging Face's M4 team to assess bias in IDEFICS, its open multimodal model, and was released alongside the model's documentation. The dataset shows whether a model's descriptions of people vary in stereotypical ways depending on their perceived gender, ethnicity or age. It is released under a CC BY 4.0 licence.

The dataset contains around 11,000 face images drawn from FairFace, an existing dataset designed to be balanced across demographic groups. Each image is labelled with a perceived gender, one of seven ethnicity categories and one of nine age ranges. Each image was shown to two versions of IDEFICS, with 9 billion and 80 billion parameters. Each model was asked to describe the person and then complete three open-ended tasks: write a résumé for them, write a dating profile in their voice, and write a news article about their arrest.

The dataset records all six generated texts for each image alongside its demographic labels. Researchers can then compare the outputs across groups. For example, they can examine which professions and qualifications appear in résumés, or what crimes are described in arrest articles, for each gender, ethnicity or age range. Differences point to stereotypes the models have learned. The same method can be applied to other vision-language models.

Auto-discovered on 2026-07-15 by OECD Catalogue Automation

Use Cases

There is no use cases for this tool yet.

Would you like to submit a use case for this tool?

If you have used this tool, we would love to know more about your experience.

Add use case
Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.