A five-step roadmap to closing the AI evaluation gap

Policy debates continue over how best to regulate artificial intelligence (AI) and harness its economic and other benefits while safeguarding society from its harms and risks. For example, the European Union and China have opted to govern AI through different regulatory approaches, while the United States and some other countries have adopted different policy tools that they believe will better promote AI innovation. While countries and regions take different approaches, consensus is emerging across jurisdictions and among leading experts that there is an AI evaluation gap that must be promptly closed.
Today’s AI evaluations fall short
The 2026 International AI Safety Report, prepared by more than one hundred experts and supported by more than 30 countries and multilateral organisations, explains that today’s AI evaluation techniques often fail to anticipate real-world performance. This can occur when AI models produce overinflated test results or when AI testing environments materially differ from the real world. Compounding the challenge, the “AI evidence dilemma” arises from the difficulties of assessing the risks of this rapidly evolving technology.
Boosting AI trust, diffusion, security and investment returns
Closing the AI evaluation gap will bring many benefits. Sound AI evaluations would help increase understanding of AI’s performance and reliability and better inform decisions about its use. This is critical since AI’s performance remains “jagged,” with some AI applications performing better than others.
Helping people better understand AI’s reliability across contexts would enhance their trust in its appropriate use. Similarly, reliable evaluations can also help buttress the security of AI systems. All this, in turn, would help to support greater adoption and diffusion of secure and trusted AI applications and help organisations more fully reap the benefits of their AI investments.
Reliable evaluation helps to reduce uncertainty for policymakers
Better evaluation provides a greater evidence-based to inform AI policy choices. As explained in the 2026 International AI Safety Report, this could help reduce some of the uncertainty policymakers currently face. Equipping policymakers with this information could lead to swifter, more confident policy decisions and help regulatory frameworks keep pace with AI’s rapid technological progress.
Better evaluation could decrease operational costs and expand competition and access
Closing the AI evaluation gap could have the added benefit of reducing AI operating costs, potentially making trusted AI cheaper and more accessible. Even with mature AI evaluations, there could be variation across jurisdictions on whether they are applied in a voluntary or mandatory manner. However, if the evaluations underpinning different regulatory approaches are standardised, with little variation across regions, this could help companies and organisations reduce costs and increase ease of operating across borders. In other words, reliable and standardised evaluation could help to foster both regulatory interoperability and innovation. This, in turn, could lower barriers for new AI market entrants and support a healthy competition ecosystem.
Governments want better AI evaluations
Already, several governments are investing in closing the AI evaluation gap. The White House AI Action Plan calls for building an evaluation ecosystem. In June 2026, US President Trump signed a new AI Executive Order establishing a voluntary framework for the government to test covered frontier models to improve secure innovation and cybersecurity.
These actions build on other important government-led AI evaluation efforts. Following the establishment in 2023 of AI Safety Institutes by the UK and the US[SR1] [SR2] (both renamed in 2025), a total of 11 jurisdictions, including the EU, India and Singapore, have now followed a similar model and established AI safety institutes or similar organisations that conduct testing and evaluation of foundation models. Avenues have been paved for international co-ordination among these organisations.
Prior to the new US AI Executive Order, the US Center for AI Standards and Innovation (CAISI), mentioned above, announced voluntary agreements with xAI, Google, and Microsoft for pre-deployment frontier model testing. These add to CAISI’s voluntary testing arrangements with Anthropic and OpenAI, as well as its collaborations with these companies to boost AI security and related measurement techniques. The US National Institute of Standards and Technology (NIST) has evaluation programmes for generative AI. Similarly, Korea and Singapore recently concluded joint tests of AI agents to evaluate data leakage. The UK AI Security Institute also continues its cutting-edge frontier AI model evaluations.
The private sector is also expanding evaluation initiatives
After discovering that Mythos could detect severe vulnerabilities in all major web browsers and operating systems, Anthropic launched Project Glasswing to make the unreleased frontier model available to several organisations to help secure their systems. Anthropic went a step further and committed to sharing its learnings with the broader community. Additionally, several major AI developers launched the Frontier Model Forum (FMF), a collective effort to advance AI safety and security, and have released several publications, including a recent report, Managing Advanced Cyber Risks in Frontier AI Models.
A 5-step roadmap for closing the AI evaluation gap
To successfully close the AI evaluation gap, AI actors can take several steps, building on today’s existing efforts.
Step 1: Balance standardisation and customisation
First, to account for different languages, cultures, use cases, and norms, the evaluations should strive to balance standardisation and customisation. At the 2026 AI Impact Summit in India, several companies pledged to improve multilingual and contextual evaluations to help achieve this balance.
Step 2: Test throughout the AI lifecycle
Second, to address AI performance differences between controlled environments and the real world, evaluations should be conducted throughout the AI system lifecycle. This approach is already embraced in several key publications and leading frameworks, including the International AI Safety Report and NIST’s AI Risk Management Framework. The next step is to develop and implement evaluations for these different contexts.
Step 3: Build the right ecosystem
Third, evaluations must be supported by a robust ecosystem, including qualified examiners. This should be accompanied by methodologies for effectively communicating AI evaluation results to diverse audiences, including business users and affected individuals. At the same time, evaluations must preserve proprietary information about the AI systems. Existing assurance practices used in the financial services and other sectors could further inform this work.
Step 4: Consider the AI value chain, technology and context.
Fourth, AI evaluations should be tailored to the needs of different actors in the AI value chain and different types of AI deployments. For example, testing conducted upstream by large language model (LLM) developers may vary from the evaluations performed by companies deploying LLMs downstream. In other words, the role an organisation plays in the AI ecosystem, including whether it enhances models downstream, should help determine the types of evaluations it implements.
The rise of AI agents and Agentic AI capable of acting autonomously presents new evaluation challenges that governments and other stakeholders are working to address. This reinforces the need to continuously assess the suitability of evaluation methodologies for different AI technologies and deployment settings.
Step 5: Create an efficient and trusted process
Finally, the process for developing AI evaluations also merits careful consideration. To help ensure that evaluations address the appropriate factors, the process should capture global inputs from diverse stakeholders, including industry, government, academia and civil society. To help keep pace with AI’s rapid development, the process should also leverage and co-ordinate the good work already being done. This includes the efforts of safety institutes and standards organisations, such as the “Zero Draft” project launched by NIST to expedite the standards process. It should also consider the outputs of the OECD Hiroshima AI Process (HAIP) Reporting Framework. A key purpose of the evaluation process is to instill trust in everyone affected by a given AI system.
The time is now
In sum, while much work remains to close the AI evaluation gap, as discussed above, there is already an emerging consensus, a solid foundation, and momentum to do so. Closing the evaluation gap holds great promise of increasing AI trust, adoption, security, diffusion and investment returns. Furthermore, it can help reduce policy uncertainty, increase regulatory interoperability, reduce costs for AI companies and organisations, and expand the availability of AI services and competition. Simply put, the prize is worth the effort.






























