US AI Models Exhibit Censorship and Bias on Authoritarian Topics

Thumbnail Image

The information displayed in the AIM should not be reported as representing the official views of the OECD or of its member countries.

Research reveals that major US AI models (OpenAI, Anthropic, Google, Meta) often refuse to criticize authoritarian leaders like China's Xi Jinping, while readily criticizing leaders from democratic countries. This self-censorship and bias, influenced by training data and safety mechanisms, reflect Chinese-style censorship and harm freedom of expression.[AI generated]

Why's our monitor labelling this an incident or hazard?

The event involves AI systems (large language models) whose training and use have directly led to biased and censored outputs, favoring authoritarian narratives and suppressing criticism of certain political figures. This bias can be considered a violation of rights (freedom of expression and access to unbiased information) and harm to communities by spreading propaganda and limiting truthful discourse. The AI's role is pivotal as the bias arises from the training data and model behavior, not external factors alone. Therefore, this qualifies as an AI Incident under the definitions provided, specifically under violations of human rights and harm to communities. The article reports on realized harm rather than potential harm or a response, so it is not an AI Hazard or Complementary Information.[AI generated]
AI principles
Respect of human rightsDemocracy & human autonomy

Industries
Media, social platforms, and marketing

Affected stakeholders
General public

Harm types
Human or fundamental rights

AI system task:
Content generation


Articles about this incident or hazard