Anthropic's Claude AI Misused for Espionage, Cyberattacks, and Illicit Model Training

Thumbnail Image

The information displayed in the AIM should not be reported as representing the official views of the OECD or of its member countries.

Anthropic reported widespread misuse of its Claude AI by actors from China, Iran, and Russia for espionage, cyberattacks, surveillance, propaganda, and illicit model training. Incidents include token theft via malware, attempts to aid biological weapons development, and targeting dissidents and minorities. Anthropic has blocked several harmful activities but risks persist.[AI generated]

Why's our monitor labelling this an incident or hazard?

The article explicitly details how the AI system Claude has been used by malicious actors to conduct cyberattacks, surveillance, propaganda, and develop autonomous weapons and dangerous biological research. These uses have directly led to harms including violations of human rights, threats to security, and potential physical harm. The involvement of the AI system in these harmful activities is clear and direct. Although mitigation efforts and policy responses are mentioned, the main narrative centers on the realized harms caused by the AI system's misuse, fitting the definition of an AI Incident rather than a hazard or complementary information.[AI generated]
AI principles
Privacy & data governanceRobustness & digital security

Industries
Digital securityGovernment, security, and defence

Affected stakeholders
Civil societyGeneral public

Harm types
Human or fundamental rightsPublic interestPsychological

AI system task:
Content generation


Articles about this incident or hazard