Catalogue of Tools & Metrics for Trustworthy AI

These tools and metrics are designed to help AI actors develop and use trustworthy AI systems and applications that respect human rights and are fair, transparent, explainable, robust, secure and safe.

Gate AI: LLM Security Benchmark Evaluation Methodology and Results



Gate AI: LLM Security Benchmark Evaluation Methodology and Results is a technical report describing a method for evaluating detectors of attacks on large language models (LLMs). These include prompt injection, where hidden instructions manipulate a model, and jailbreaks, which bypass safety controls. Published evaluations of such detectors are often hard to compare. They tend to use different datasets, adjust detection thresholds separately for each benchmark, and do not always disclose the settings used. The report sets out a more consistent approach and applies it to Gate AI, a commercial security gateway for AI applications. It was published in June 2026 by Constellation Network as a working preprint that will be updated over time.

The detector is tested on 16 public benchmarks containing 12 111 samples, using five-fold cross-validation. This means it is repeatedly tested on data held back from training. A second analysis groups near-duplicate prompts together, so that near-identical examples cannot appear in both training and test data and inflate results. A single detection threshold is chosen, allowing at most 1% false alarms, and applied to every benchmark. Comparisons with competing detectors then adjust thresholds so all systems are compared at the same false-alarm rate.

Further checks test how well the detector generalises. These include leaving one benchmark out of training entirely and training on deliberately randomised labels. The authors report that Gate AI detects about 95% of attacks at this false-alarm limit. The report also acknowledges gaps, such as attacks spread across multiple conversation turns.

Auto-discovered on 2026-07-15 by OECD Catalogue Automation

Use Cases

There is no use cases for this tool yet.

Would you like to submit a use case for this tool?

If you have used this tool, we would love to know more about your experience.

Add use case
Partnership on AI

Disclaimer: The tools and metrics featured herein are solely those of the originating authors and are not vetted or endorsed by the OECD or its member countries. The Organisation cannot be held responsible for possible issues resulting from the posting of links to third parties' tools and metrics on this catalogue. More on the methodology can be found at https://oecd.ai/catalogue/faq.