There is an uncomfortable result in medical AI that keeps getting replicated, and it is not the one you have heard about.
The headline version is familiar by now: a language model beats physicians on diagnostic vignettes. Fine. Vignettes are not patients, and most of us have learned to discount that.
The stranger finding is this one. When you give physicians the same model and let them use it, they score worse than the model did on its own. Not worse than they would have done without it. Worse than the machine alone.
A team in London just replicated that result and then did the thing nobody had bothered to do: they read the interaction logs to find out why. The answer is not that doctors are stubborn, or that the model is bad. It is considerably more interesting than either.
It turns out the limiting factor was not the model, and it was not the physicians either. It was the space between them.
Let's get into it.
— Troy, Ray, and Ibrahim
The AI scored higher alone than it did with a doctor attached
Physicians at University College Hospital in London set out to replicate a widely discussed 2024 JAMA Network Open trial by Goh and colleagues, which found a large language model working alone outperformed American clinicians who had access to that same model. The UK team ran a within-subjects study: 22 physicians, four clinical vignettes, model access on two of them, analyzed with mixed-effects models accounting for case difficulty and baseline clinician variability.
The replication held. Physicians with model assistance scored significantly lower than the model alone, a gap of 21.3 percentage points (P<.001). Then the researchers coded the interaction logs.
Why it matters: Two things were true at once, and the second one rarely makes the headline. Access to the model did improve physician performance compared with conventional resources, 74.3% versus 65.7% (P=.001). The tool helped. It just helped far less than it could have. And the logs showed why: physicians actually posed only about 30% of the case questions to the model. The performance gap was substantially a usage gap.
The catch: Twenty-two physicians, four vignettes, no real patients, no time pressure, no consequences for being wrong. The improvement was also strikingly uneven between individual clinicians (standard deviation 12.8%), which means the average conceals people who got a lot out of the tool and people who got very little. And a replication of an artificial task is still an artificial task.
Bottom line: Stop reading these studies as a scoreboard between humans and machines. The variable that moved performance was not intelligence on either side. It was whether the clinician actually asked. That is a training problem and an interface problem, and both are fixable in a way that neither party's raw ability is.
Read the full report from Diagnosis →
Where do you land on this one? Hit reply and tell us. We read every response, and the best ones shape the next issue.
From the Field
Here is our question this week, and we suspect the answers will be honest in a way conference talks are not. When you have a model open in another tab, do you actually use it? On what kinds of problems, and at what point in your reasoning? And the harder version: have you ever not asked because you did not want to find out you were wrong? Attendings, residents, students, nurses, anyone. We protect your anonymity and your patients' privacy, always.
The Profile: Adam Rodman, MD, MPH, FACP
General internist and Director of AI Programs, Beth Israel Deaconess Medical Center; Assistant Professor, Harvard Medical School
Adam Rodman is a general internist, a medical historian, and one of the few researchers treating clinical reasoning as the actual object of study rather than a backdrop for benchmarking models.
His work keeps asking the question this week's feature raises: not whether the machine is smart, but what happens to a physician's thinking when the machine is in the room. He also leads the integration of AI into the medical curriculum at BIDMC, which means he is building the training layer the London study says is missing.
- Corresponding author on a 2023 JAMA Network Open study comparing a chatbot's probabilistic reasoning against practicing clinicians
- Co-authored a 2025 Nature Medicine randomized trial finding physicians with large language model access performed better on patient care tasks than those without
- Serves as Director of AI Programs at Beth Israel Deaconess Medical Center, leading AI integration into medical education
- Assistant Professor at Harvard Medical School, with research spanning clinical reasoning, medical history, and human-computer interaction
Read more at Beth Israel Deaconess Medical Center →
Quick Hits
A language model held its own in complex cardiology
In a randomized trial published in Nature Medicine, nine general cardiologists managed complex suspected genetic cardiomyopathy cases with or without an experimental Google model. Blinded subspecialists preferred the assisted assessments 46.7% of the time versus 32.7% for cardiologists alone (P=.02). Unassisted cardiologists had more clinically significant errors, 24.3% versus 13.1%. LEARN MORE
The original result that started the argument
Goh and colleagues reported in JAMA Network Open in 2024 that a large language model working alone outperformed physicians who had access to that same model. It has been cited relentlessly and misread almost as often. The UK replication this week is the first to explain the mechanism rather than just repeat the finding. LEARN MORE
The other failure mode: believing it too easily
Under-use is one risk. Automation bias is the other. Researchers have been running trials in which physicians trained in AI literacy receive model suggestions seeded with deliberate, clinically significant errors, to measure how often the errors survive review. Under-asking and over-trusting are opposite problems with the same root: no calibrated sense of when the tool is reliable. LEARN MORE
The FDA quietly redrew the line in January
Revised FDA guidance issued January 6, 2026 superseded the 2022 clinical decision support guidance, clarifying which functions stay outside device regulation under the 21st Century Cures Act. Software offering a single recommendation for provider review may remain unregulated. Anything analyzing medical images for diagnosis does not. LEARN MORE
Off the Clock: the skill nobody is teaching
Thirty percent. That is the number worth carrying out of this issue.
Given a model that demonstrably knew the answers, experienced physicians consulted it on fewer than one in three questions. Not because they distrusted it, exactly. The researchers did not find refusal. They found something closer to forgetting it was there.
Medicine has a long history with this. The literature on clinical decision support is thirty years of the same discovery: the tool is accurate, the tool is available, and the tool is not used at the moment it would have mattered. We built alerts, and people clicked through them. We built order sets, and people worked around them. Now we have built something that can genuinely reason, and the early evidence says we will mostly forget to ask it.
There is an old apprenticeship idea buried in here. Knowing what you do not know, and knowing precisely when to turn to someone who does, has always been the core competency of a good clinician. That is what a curbside consult is. It is what the attending on the other end of the phone at 3am is for. Nobody ever called that skill outsourcing.
What is new is that the consultant is now available for every patient, instantly, at no social cost, and with no risk of looking uncertain in front of a colleague. You would think that would increase asking. The data says it has not, yet.
Which suggests the barrier was never access. It was the habit of recognizing the moment. That habit was built in us by people who asked, out loud, in front of us, when they did not know something.
The machine cannot teach that. The attending on the floor still can.
Until next Sunday.
For the people keeping medicine human.
