Persistent AI Hallucinations Highlight Risks in Critical Applications

Thumbnail Image

The information displayed in the AIM should not be reported as representing the official views of the OECD or of its member countries.

Recent research and expert warnings highlight that hallucinations—false outputs generated by large language models (LLMs)—are unavoidable and increase with input size. These inaccuracies pose significant risks in high-stakes fields like law and accounting, challenging the reliability of AI for critical tasks.[AI generated]

Why's our monitor labelling this an incident or hazard?

The event involves AI systems (LLMs) and their use, specifically their tendency to hallucinate false outputs. Although no direct harm is described as having occurred, the article clearly outlines the potential for these hallucinations to cause significant harm in critical domains. Therefore, this situation fits the definition of an AI Hazard, as the development and use of these AI systems could plausibly lead to an AI Incident involving harm to persons, organizations, or communities relying on accurate outputs.[AI generated]
AI principles
Robustness & digital securitySafety

Industries
Financial and insurance servicesGovernment, security, and defence

Affected stakeholders
ConsumersBusiness

Harm types
Economic/PropertyReputational

Business function:
Compliance and justice

AI system task:
Content generation


Articles about this incident or hazard