Flaw in Major AI APIs Exposes Hidden Reasoning and Sensitive Data

Thumbnail Image

The information displayed in the AIM should not be reported as representing the official views of the OECD or of its member countries.

Researchers discovered a vulnerability in the APIs of OpenAI, Anthropic, and Google that allowed weaker AI models to extract hidden reasoning steps, API keys, and passwords from encrypted reasoning blocks. Thousands of private artifacts were exposed from public logs, revealing significant privacy and security risks due to flawed AI system design.[AI generated]

Why's our monitor labelling this an incident or hazard?

The article explicitly mentions a flaw in AI systems from major providers (OpenAI, Anthropic, Google) that allowed recovery of internal reasoning and secrets, including API keys and passwords. This is a direct malfunction of AI systems leading to harm through exposure of sensitive data, which is a violation of intellectual property and privacy rights. The involvement of AI systems is clear, and the harm is realized, not just potential. Hence, this qualifies as an AI Incident.[AI generated]
AI principles
Privacy & data governanceRobustness & digital security

Industries
Digital security

Affected stakeholders
ConsumersBusiness

Harm types
Human or fundamental rightsReputational

AI system task:
Content generationReasoning with knowledge structures/planning


Articles about this incident or hazard