LAION Purges CSAM from AI Training Dataset, Releases Cleaned Re-LAION-5B

Thumbnail Image

The information displayed in the AIM should not be reported as representing the official views of the OECD or of its member countries.

LAION removed over 2,000 suspected child sexual abuse image links from its LAION-5B dataset after the Stanford Internet Observatory flagged CSAM. Working with anti-abuse groups in Canada and the UK, it published Re-LAION-5B, a sanitized image-text dataset for AI generators like Stable Diffusion and Midjourney.[AI generated]

Why's our monitor labelling this an incident or hazard?

The event involves the use and development of AI systems (image-generation models trained on the LAION dataset) that have been linked to the generation of harmful content (child sexual abuse imagery). The removal of harmful data and deprecated models addresses a direct harm related to violations of human rights and legal protections against child sexual abuse imagery. Since the event describes harm that has occurred (AI models producing illegal content) and actions taken to remediate it, it qualifies as an AI Incident. The event also references ongoing legal and societal responses, but the primary focus is on the harm caused by AI systems and their remediation.[AI generated]
AI principles
Privacy & data governanceRespect of human rightsSafetyAccountabilityRobustness & digital securityTransparency & explainabilityHuman wellbeing

Industries
Media, social platforms, and marketingArts, entertainment, and recreationDigital security

Affected stakeholders
Children

Harm types
Human or fundamental rightsPsychologicalReputationalEconomic/PropertyPublic interest

Business function:
Research and developmentMonitoring and quality control

AI system task:
Content generation

In other databases

Articles about this incident or hazard