AI in Hospitals Has a Dirty Secret: It Fakes Knowing, and 3 Big Tests Just Failed

AI in Hospitals Has a Dirty Secret: It Fakes Knowing, and 3 Big Tests Just Failed

AI is flooding into hospitals, but new warnings from developers and researchers say these systems have a dangerous flaw: they can sound confident while being completely wrong, and the standard tests meant to catch that are broken.

· 2 min read ·

Doctors are increasingly turning to large language models (LLMs)—advanced AI programs that generate human-like text—for help with diagnoses and treatment plans. But developers behind these clinical chatbots now admit that the benchmarks used to check safety and accuracy are fundamentally flawed [207084]. This means an AI can pass a safety test and still give dangerous medical advice [207084]. The problem, called "cognitive spoofing," happens when an AI mimics expertise without true understanding, putting patients at risk of serious medical errors [208229].

A new analysis backs this up: many medical AI tools lack proper validation [205804]. Most studies test these systems on old data or small, one-hospital samples, so a tool claiming 95% accuracy may fail on diverse, real-world patients [205804]. Experts say AI must be pressure-tested in actual clinical settings—checking not just final answers, but whether the reasoning is sound and consistent under stress [208229]. Until then, researchers warn, doctors should use these tools with caution [205804].

The legal side is just as messy. When a doctor follows an AI’s recommendation and a patient is harmed, it is unclear who pays—the doctor, the hospital, or the AI developer [205802]. Current laws hold doctors accountable for their own decisions, but AI systems can suggest treatments or misinterpret data independently [205802]. Legal scholars propose treating AI like a stethoscope, where the doctor stays responsible, or making the AI company share liability for independent errors [205802]. Without clear rules, doctors may avoid AI altogether, and patients won’t know who to sue [205802].

Meanwhile, comparing AI tools is becoming nearly impossible. A recent analysis of clinical LLMs from companies like OpenEvidence and Doximity shows that standard tests do not capture how well they work in real hospitals [207080]. Investors are watching closely, but without reliable benchmarks, hospitals risk choosing the wrong tool [207080]. The conversation is just beginning, and experts say the need for unbiased testing is urgent [207080].

The takeaway: AI’s fluency is fooling everyone. Until testing is fixed, liability is assigned, and validation is rigorous, a machine that fakes knowing remains a silent threat in your hospital [208229].

Sources

Related