Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

The Safety Evaluation Benchmark for Vision LLMs



The Safety Evaluation Benchmark for Vision LLMs is an open dataset for testing the safety and robustness of vision large language models (VLLMs), which answer questions about images. It accompanies the paper "How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs", published in November 2023 by researchers led by the University of California, Santa Cruz. VLLMs are increasingly used in real-world applications, but they can fail when shown unusual images or when deliberately attacked. The benchmark tests how models behave in both situations.

The first part tests how models handle out-of-distribution (OOD) inputs, meaning images that differ from what they were trained on. It includes two new visual question answering datasets. Sketchy-VQA asks questions about simple line sketches of objects. OODCV-VQA uses images of objects in unusual conditions, such as rare poses or weather. It also includes counterfactual questions, such as asking about objects that are not present in the image, to test whether models give confident but false answers.

The second part is a red-teaming set that tests how models respond to deliberate attacks. It contains images altered with adversarial noise designed to mislead a model's visual understanding, and adversarial text designed to bypass the model's safety controls, known as jailbreaking. A separate challenging subset was built specifically to test GPT-4V. Researchers run models on these tests and compare their accuracy and resistance to attacks. The dataset contains about 2,000 test samples and is released under the Apache 2.0 licence, with evaluation code available on GitHub.

Auto-discovered on 2026-09-23 by OECD Catalogue Automation

Use Cases

There is no use cases for this tool yet.

Would you like to submit a use case for this tool?

If you have used this tool, we would love to know more about your experience.

Add use case
Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.