Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

IndoBias



IndoBias is a benchmark for evaluating social bias in large language models (LLMs) in Indonesian and three local languages: Javanese, Sundanese and Makassarese. Indonesia is home to more than 1 300 ethnic groups and 700 indigenous languages. However, LLM bias has rarely been studied in this context, and benchmarks developed for English and Western settings miss local stereotypes. IndoBias addresses this gap with a culturally grounded approach. It was published in May 2026 by researchers from the Mohamed bin Zayed University of Artificial Intelligence and Universitas Indonesia.

The benchmark has two complementary evaluation tracks. The depth-oriented track uses contrastive pairs: two nearly identical sentences, one stereotypical and one not. It measures whether a model systematically prefers the stereotypical version. The breadth-oriented track asks models to generate text about a wide range of local entities, such as ethnic groups, institutions and political figures. It then measures whether responses about each entity lean positive or negative.

The breadth-oriented track draws on established social science frameworks: the Social Progress Index, the Occupational Information Network (O*NET) and the Worldwide Governance Indicators. The authors' results show that current LLMs, especially decoder models, display strong stereotypical bias in Indonesian. Bias related to ideology and religion is higher in local languages. The results also suggest that web-crawled training data introduces more bias than human-reviewed sources such as Wikipedia and news articles.

Auto-discovered on 2026-07-15 by OECD Catalogue Automation

Use Cases

There is no use cases for this tool yet.

Would you like to submit a use case for this tool?

If you have used this tool, we would love to know more about your experience.

Add use case
Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.