
The information displayed in the AIM should not be reported as representing the official views of the OECD or of its member countries.
The Wikimedia Foundation has raised alarms over AI training methods that use automated web crawlers to extract vast amounts of data, overwhelming Wikipedia and related servers. This excessive data scraping is causing rising operational costs and poses risks of service disruptions, highlighting a growing digital infrastructure hazard.[AI generated]
Why's our monitor labelling this an incident or hazard?
The event involves AI systems (large language models) that require massive data scraping from Wikimedia sites via automated bots (crawlers). This automated AI-driven data extraction has directly led to operational disruptions and increased costs for Wikimedia, which is a form of harm to infrastructure and community resources. The harm is realized, not just potential, as Wikimedia reports slowdowns and resource strain. Hence, it meets the criteria for an AI Incident because the AI system's use (data scraping for training) has indirectly caused harm (disruption and resource strain) to a critical information infrastructure.[AI generated]