
The information displayed in the AIM should not be reported as representing the official views of the OECD or of its member countries.
A study by National Taiwan University evaluated ChatGPT, Claude, and Gemini in providing stroke care information, finding all three generative AI models performed below clinical safety thresholds, often giving incomplete or incorrect advice. Researchers warn that relying on these AI systems as medical authorities could endanger patient health.[AI generated]
Why's our monitor labelling this an incident or hazard?
The article explicitly involves AI systems (generative large language models) used to provide medical information to stroke patients. The research identifies significant accuracy and reliability issues that could lead to fatal harm if patients depend on AI advice for critical medical decisions. No actual harm event is reported, but the credible risk of fatal outcomes due to AI misinformation in a high-stakes medical context meets the definition of an AI Hazard. The article does not describe a realized incident but warns of plausible future harm, thus it is not an AI Incident. It is not merely complementary information because the main focus is on the risk assessment of AI use in stroke care, not on responses or governance. Therefore, the classification is AI Hazard.[AI generated]