Anthropic AI Model Source Code Leak and Restricted Release Due to Security Risks

Thumbnail Image

The information displayed in the AIM should not be reported as representing the official views of the OECD or of its member countries.

Anthropic accidentally leaked the source code of its Claude Code AI system, exposing proprietary information but not client data. Separately, Anthropic restricted access to its powerful new AI model, Claude Mythos Preview, due to its unprecedented ability to identify software vulnerabilities, fearing misuse by malicious actors and potential cybersecurity threats.[AI generated]

Why's our monitor labelling this an incident or hazard?

The article explicitly mentions an AI system (Claude Mythos Preview) with advanced capabilities in cybersecurity vulnerability detection. Anthropic limits access to prevent malicious exploitation, indicating awareness of potential misuse risks. No direct or indirect harm has yet occurred, but the model's power and potential for misuse pose a credible risk of harm to critical infrastructure and security. Hence, this event fits the definition of an AI Hazard, as the AI system's development and potential use could plausibly lead to an AI Incident in the future.[AI generated]
AI principles
Robustness & digital security

Industries
Digital security

Affected stakeholders
Business

Harm types
Economic/PropertyReputational

Business function:
Research and development

AI system task:
Content generationEvent/anomaly detection


Articles about this incident or hazard