Researchers Use Anthropic's Claude AI to Breach OpenAI Systems

Thumbnail Image

The information displayed in the AIM should not be reported as representing the official views of the OECD or of its member countries.

A team of cybersecurity researchers from Hacktron AI used Anthropic's Claude AI to exploit vulnerabilities and gain unauthorized access to OpenAI employee ChatGPT accounts and a private GitHub repository. The breach, conducted as part of a bug bounty program, demonstrated AI's role in accelerating exploit development. OpenAI awarded a $6,500 bounty.[AI generated]

Why's our monitor labelling this an incident or hazard?

The AI system (Claude) was actively used to hack into OpenAI's systems, leading to unauthorized access to private accounts and repositories. This is a direct AI Incident because the AI's use directly led to a security breach and potential harm to property and privacy rights. Although the hack was reported through an authorized bug bounty program and the researchers were rewarded, the event still qualifies as an AI Incident due to the realized harm from the AI-enabled exploit.[AI generated]
AI principles
Privacy & data governanceRobustness & digital security

Industries
Digital security

Affected stakeholders
WorkersBusiness

Harm types
Human or fundamental rights

Business function:
ICT management and information security

AI system task:
Content generationReasoning with knowledge structures/planning


Articles about this incident or hazard