Somewhere in your hospital right now, a model is firing an alert about a patient you have not seen yet.
Most of us have stopped asking where it came from. It is in the EHR, it has a name, it has been there for years, and the assumption is that something this consequential must have been checked by someone.
This Sunday we want to revisit the study that proved that assumption wrong, because it remains the single most useful thing a working clinician can know about clinical AI. Not because the story is new. Because the conditions that produced it have not meaningfully changed, and the tools arriving on your floor this year are being adopted on exactly the same faith.
The short version: a proprietary sepsis model deployed at hundreds of US hospitals turned out, on independent examination, to miss roughly two thirds of the patients it existed to catch.
Let's get into it.
— Troy, Ray, and Ibrahim
The sepsis model hundreds of hospitals trusted missed two thirds of cases
In 2021, a team at the University of Michigan published an external validation of the Epic Sepsis Model in JAMA Internal Medicine. The model was, at the time, implemented at hundreds of US hospitals. It had never been adequately evaluated by anyone outside the company that built it.
The researchers studied 27,697 adult patients across 38,455 hospitalizations at Michigan Medicine between December 2018 and October 2019, recalculating the model's score every 15 minutes. Sepsis occurred in 2,552 hospitalizations, about 7%.
Why it matters: The model's area under the ROC curve was 0.63 (95% CI 0.62 to 0.64). It failed to identify 1,709 of the 2,552 patients with sepsis, about 67%. It did this while generating alerts on 6,971 of all hospitalizations, roughly 18% of every patient who came through the doors. Miss most of the real cases, interrupt nearly one in five clinicians anyway. That is the worst possible combination: low sensitivity paired with a heavy alert burden, which is a machine for manufacturing alert fatigue.
The catch: The catch is not with the study. It is with the fact that this took until 2021, and that the pattern persists. A 2024 external validation in two county emergency departments, published in JAMIA Open, found a sensitivity of 14.7% within a six-hour window. Version 2 of the model has still not had a published multicenter independent validation. And proprietary models remain difficult to scrutinize precisely because they are proprietary.
Bottom line: Ask one question of every predictive model running in your shop: who validated this, on whose patients, and was it anyone other than the people selling it? If nobody can answer, you are not using a clinical tool. You are using a hypothesis with a deployment contract.
Read the full report from JAMA Internal Medicine →
Where do you land on this one? Hit reply and tell us. We read every response, and the best ones shape the next issue.
From the Field
We want the alert stories this week. The predictive model everyone on your unit learned to dismiss, and what it cost the one time it was right. The rollout that arrived with training slides and no validation data. The nurse or resident who noticed the thing was wrong months before anyone with authority did. These are the stories that never make it into the literature and should. We protect your anonymity and your patients' privacy, always.
The Profile: Karandeep Singh, MD, MMSc
Inaugural Chief Health AI Officer, UC San Diego Health; Joan and Irwin Jacobs Endowed Chair in Digital Health Innovation
Karandeep Singh was the senior author on the Michigan study that took apart the Epic Sepsis Model. He is a nephrologist and informatician, and his central argument has been consistent for years: a model deployed in patient care is a clinical intervention, and clinical interventions get evaluated.
He now holds the job that argument implies. UC San Diego Health created the role of Chief Health AI Officer and hired him into it, which is roughly what it looks like when a health system decides that scrutiny should be somebody's actual title rather than everybody's spare-time concern.
- Senior author of the 2021 JAMA Internal Medicine external validation that found the Epic Sepsis Model missed 67% of sepsis cases while alerting on 18% of all hospitalizations
- Named inaugural Chief Health AI Officer at UC San Diego Health, effective December 2023
- Holds the Joan and Irwin Jacobs Endowed Chair in Digital Health Innovation at UC San Diego School of Medicine
- Oversees clinical deployment and evaluation of predictive models at UCSD, including sepsis early-warning work in the emergency departments
Read more at UC San Diego Health →
Quick Hits
A second validation, even worse numbers
Researchers externally validated the Epic sepsis predictive model in two county emergency departments and published the results in JAMIA Open. Within a six-hour window, sensitivity was 14.7%, with a positive predictive value of 7.6%. Different setting, same conclusion, three years later. LEARN MORE
Certification arrives for AI governance
The Joint Commission launched its voluntary Responsible Use of AI in Healthcare certification in June 2026, covering governance, safeguards, monitoring, and education. It builds on guidance released with the Coalition for Health AI in September 2025. Voluntary is doing a lot of work in that sentence, but it is a structure where there was none. LEARN MORE
The case for assurance labs
The Coalition for Health AI has been building toward a national network of certified labs that would test health AI products for safety and effectiveness, along with model cards functioning as nutrition labels. The premise is simple and overdue: independent evaluation should not depend on an academic center deciding to do it for free. LEARN MORE
A maturity model for health AI governance
A systematic review in npj Digital Medicine proposes a comprehensive maturity model for healthcare AI governance, giving organizations a way to assess where they actually sit rather than whether they have a committee. Useful reading for anyone whose job title recently acquired the letters AI. LEARN MORE
Off the Clock: the alarm that cried wolf
Alert fatigue is usually described as a human failing. Clinicians become desensitized. Clinicians click through. Clinicians stop reading.
The Michigan numbers suggest a less flattering account of who failed whom.
A model that fires on 18% of all hospitalizations while catching a third of the actual cases is not a warning system that clinicians unfortunately learned to ignore. It is a warning system that taught them to ignore it, correctly, through repetition. The desensitization was not a lapse in vigilance. It was accurate learning from the available evidence.
There is a version of the boy who cried wolf that nobody tells, where the villagers are the ones held responsible for not responding to the ninety-ninth false alarm. That is roughly the position bedside staff have been placed in for a decade.
What makes this worth revisiting in 2026 is that the mechanism is about to repeat at a much larger scale. Ambient tools, agentic workflows, models drafting messages and flagging risk and queueing orders. Most of them will arrive the same way the sepsis model did: procured, integrated, announced, and never independently examined against your patients. The staff will be trained on how to use them. Almost nobody will be shown the validation data, because in many cases it will not exist.
And the people who will notice first, as always, are the ones closest to the patient. The nurse who realizes the score is high on everyone. The resident who stops trusting the flag. Their skepticism will be described as resistance to change, when it is in fact the only functioning quality control in the building.
If your hospital wants those people to keep trusting the tools, there is one reliable way to earn it. Show them the numbers. On their patients. From someone who was not paid to produce them.
Until next Sunday.
For the people keeping medicine human.
