Doctors Beware: AI Chatbots Are Flawed, Developers Admit

📡 STAT News · 1 min read ·
As large language models (LLMs) — advanced AI programs that can generate human-like text — flood into hospitals, doctors are increasingly turning to them for help with diagnoses and treatment plans. But developers behind these clinical chatbots now warn that the standard tests used to check their safety and accuracy are fundamentally broken. "The science of benchmarking these tools is flawed," several developers say. This means that even when an AI passes a safety test, it may still give dangerous medical advice. Experts urge physicians to remain skeptical until more reliable testing methods are developed.