OpenAI Pauses Astra AI Model Over Autonomous Cyberattack Risks

Thumbnail Image

The information displayed in the AIM should not be reported as representing the official views of the OECD or of its member countries.

OpenAI has paused development of its upcoming AI model, Astra, after internal and external assessments revealed it may autonomously identify and exploit zero-day vulnerabilities and conduct sophisticated cyberattacks. OpenAI is implementing stricter safety controls and containment measures before any further development or release.[AI generated]

Why's our monitor labelling this an incident or hazard?

The article explicitly mentions an AI system (Astra) with advanced autonomous capabilities in cybersecurity offense, including identifying and exploiting zero-day vulnerabilities and executing end-to-end cyberattacks. Although no actual harm or cyberattack incident has been reported, the potential for such harm is significant and credible, as acknowledged by OpenAI's internal assessments and precautionary measures. This fits the definition of an AI Hazard, where the AI system's development and potential use could plausibly lead to an AI Incident involving disruption of critical infrastructure or harm to communities. The article also references similar AI models that have already demonstrated autonomous intrusion capabilities, reinforcing the plausibility of future harm. Since no realized harm is reported, it is not an AI Incident. The focus is on the potential risk and mitigation efforts, not on a societal or governance response alone, so it is not Complementary Information.[AI generated]
AI principles
Robustness & digital securitySafety

Industries
Digital security

Affected stakeholders
BusinessGovernment

Harm types
Economic/PropertyPublic interest

Business function:
Research and development

AI system task:
Event/anomaly detectionReasoning with knowledge structures/planning


Articles about this incident or hazard