AI in Hospitals Has a Dirty Secret: It Fakes Knowing, and 3 Big Tests Just Failed
AI is flooding into hospitals, but new warnings from developers and researchers say these systems have a dangerous flaw: they can sound confident while being completely wrong, and the standard tests meant to catch that are broken.
Doctors are increasingly turning to large language models (LLMs)—advanced AI programs that generate human-like text—for help with diagnoses and treatment plans. But developers behind these clinical chatbots now admit that the benchmarks used to check safety and accuracy are fundamentally flawed [207084]. This means an AI can pass a safety test and still give dangerous medical advice [207084]. The problem, called "cognitive spoofing," happens when an AI mimics expertise without true understanding, putting patients at risk of serious medical errors [208229].
A new analysis backs this up: many medical AI tools lack proper validation [205804]. Most studies test these systems on old data or small, one-hospital samples, so a tool claiming 95% accuracy may fail on diverse, real-world patients [205804]. Experts say AI must be pressure-tested in actual clinical settings—checking not just final answers, but whether the reasoning is sound and consistent under stress [208229]. Until then, researchers warn, doctors should use these tools with caution [205804].
The legal side is just as messy. When a doctor follows an AI’s recommendation and a patient is harmed, it is unclear who pays—the doctor, the hospital, or the AI developer [205802]. Current laws hold doctors accountable for their own decisions, but AI systems can suggest treatments or misinterpret data independently [205802]. Legal scholars propose treating AI like a stethoscope, where the doctor stays responsible, or making the AI company share liability for independent errors [205802]. Without clear rules, doctors may avoid AI altogether, and patients won’t know who to sue [205802].
Meanwhile, comparing AI tools is becoming nearly impossible. A recent analysis of clinical LLMs from companies like OpenEvidence and Doximity shows that standard tests do not capture how well they work in real hospitals [207080]. Investors are watching closely, but without reliable benchmarks, hospitals risk choosing the wrong tool [207080]. The conversation is just beginning, and experts say the need for unbiased testing is urgent [207080].
The takeaway: AI’s fluency is fooling everyone. Until testing is fixed, liability is assigned, and validation is rigorous, a machine that fakes knowing remains a silent threat in your hospital [208229].