Experts Warn of AI Safety Risks from OpenAI's Astra Model's Hidden Reasoning

Thumbnail Image

The information displayed in the AIM should not be reported as representing the official views of the OECD or of its member countries.

AI safety experts have raised concerns about OpenAI's upcoming Astra model, which uses a technique called "recurrent depth" or "opaque recurrence." This approach makes the model's internal reasoning less transparent, potentially undermining safety monitoring and increasing the risk of undetected harmful AI behavior. Similar techniques are reportedly considered by Anthropic and Google DeepMind.[AI generated]

Why's our monitor labelling this an incident or hazard?

The event involves the development and use of an AI system (OpenAI's Astra model) employing a novel technique that obscures its internal reasoning from auditors, thereby undermining existing safety monitoring methods. While no actual harm has been reported, the article highlights credible expert concerns that this opacity could lead to undetectable misbehavior or unsafe AI actions in the future, especially as other major labs may adopt the same approach, potentially triggering a race to the bottom in transparency. This plausible future risk of harm from the AI system's use fits the definition of an AI Hazard. It is not an AI Incident because no realized harm has occurred yet, nor is it Complementary Information since the main focus is on the risk posed by the technique itself rather than a response or update to a past incident.[AI generated]
AI principles
SafetyTransparency & explainability

Industries
IT infrastructure and hostingDigital security

Affected stakeholders
General public

Harm types
Other

Business function:
Research and development

AI system task:
Content generationReasoning with knowledge structures/planning


Articles about this incident or hazard