North Korea-Linked Gaslight Malware Uses Prompt Injection to Evade AI Analysis

Thumbnail Image

The information displayed in the AIM should not be reported as representing the official views of the OECD or of its member countries.

Researchers discovered macOS.Gaslight, a North Korea-linked malware that employs prompt injection attacks to deceive AI-assisted malware analysis tools. By embedding fabricated system messages, the malware manipulates large language model-based triage agents, causing them to misinterpret or abort analysis, enabling data theft and system compromise. The incident highlights AI system vulnerabilities in cybersecurity.[AI generated]

Why's our monitor labelling this an incident or hazard?

The event involves an AI system explicitly, namely AI-assisted triage tools and LLMs used in malware analysis. The malware's prompt injection is designed to manipulate these AI systems, causing them to stop analyzing the malware, which indirectly leads to harm by enabling the malware to operate undetected. This constitutes an AI Incident because the AI system's malfunction (being misled by adversarial input) directly contributes to harm (security breaches, data theft). The article reports an actual discovered malware exploiting AI systems, not just a potential risk, so it is not merely a hazard or complementary information.[AI generated]
AI principles
Robustness & digital securitySafety

Industries
Digital security

Affected stakeholders
Business

Harm types
Economic/PropertyHuman or fundamental rights

Business function:
ICT management and information security

AI system task:
Event/anomaly detection


Articles about this incident or hazard