Twitter's Automated Moderation Linked to Surge in Harmful Content

Thumbnail Image

The information displayed in the AIM should not be reported as representing the official views of the OECD or of its member countries.

Following Elon Musk's acquisition, Twitter shifted to AI-driven automated moderation, reducing manual reviews and favoring content visibility restrictions over removals. This approach coincided with a reported surge in hate speech and child exploitation material, raising concerns about the effectiveness and unintended consequences of AI moderation on user safety and benign content.[AI generated]

Why's our monitor labelling this an incident or hazard?

The article explicitly mentions that Twitter replaced much of its human moderation staff with automated systems for content moderation. This AI system's malfunction or inadequate performance has directly led to increased hate speech and harmful content on the platform, causing harm to communities and violating rights. The harm is realized and ongoing, meeting the criteria for an AI Incident. The AI system's role is pivotal as it automates moderation decisions that have resulted in increased harmful content.[AI generated]
AI principles
SafetyRespect of human rightsAccountabilityTransparency & explainabilityRobustness & digital securityFairnessDemocracy & human autonomyHuman wellbeing

Industries
Media, social platforms, and marketingDigital security

Affected stakeholders
General publicChildren

Harm types
PsychologicalHuman or fundamental rightsPublic interestReputational

Business function:
Monitoring and quality controlICT management and information security

AI system task:
Recognition/object detectionOrganisation/recommendersEvent/anomaly detectionOther


Articles about this incident or hazard