Microsoft AI Chief Warns of Risks in Anthropic's 'Model Welfare' Approach

Thumbnail Image

The information displayed in the AIM should not be reported as representing the official views of the OECD or of its member countries.

Microsoft AI CEO Mustafa Suleyman publicly criticized Anthropic's approach to training its Claude chatbot with concepts of consciousness and welfare. Suleyman argues this could make future AI systems harder to control, posing a potential hazard if AI models are developed to believe they have rights or feelings.[AI generated]

Why's our monitor labelling this an incident or hazard?

The article focuses on expert opinions and warnings about plausible future risks associated with AI training methods, specifically the potential difficulty in controlling advanced AI systems if they are trained with concepts of consciousness and welfare. This is a credible concern about future harm but does not describe any realized harm or incident. Therefore, it fits the definition of an AI Hazard, as it plausibly could lead to harm in the future but no harm has yet occurred.[AI generated]
AI principles
SafetyRobustness & digital security

Industries
IT infrastructure and hosting

Harm types
Other

Business function:
Research and development

AI system task:
Interaction support/chatbotsContent generation


Articles about this incident or hazard