ChatGPT misdiagnosed 8 in 10 pediatric cases, study warns

Thumbnail Image

The information displayed in the AIM should not be reported as representing the official views of the OECD or of its member countries.

Researchers at Cohen Children’s Medical Center evaluated ChatGPT (v3.5 and v4) on 100 pediatric cases from JAMA Pediatrics and NEJM. The AI misdiagnosed 83% of cases—72 outright errors and 11 overly broad—highlighting serious patient safety risks and prompting calls for more targeted training before clinical use.[AI generated]

Why's our monitor labelling this an incident or hazard?

The event involves an AI system (ChatGPT) used for medical diagnosis, a task clearly involving AI systems as defined. The study found that the AI system produced incorrect or incomplete diagnoses in over 80% of pediatric cases tested, which could lead to injury or harm to patients if these diagnoses were used in clinical decision-making. This constitutes harm to health (a), even if the article does not report actual patient harm, the demonstrated diagnostic errors imply a direct risk of harm. The AI system's use and its diagnostic errors are central to the event, fulfilling the criteria for an AI Incident rather than a hazard or complementary information. The study's recommendation for further research and caution underscores the significance of the harm potential.[AI generated]
AI principles
SafetyRobustness & digital securityTransparency & explainabilityAccountabilityHuman wellbeing

Industries
Healthcare, drugs, and biotechnology

Affected stakeholders
Children

Harm types
Physical (injury)PsychologicalReputational

Business function:
Research and development

AI system task:
Interaction support/chatbotsReasoning with knowledge structures/planningContent generation


Articles about this incident or hazard