The Opening
Some weeks, medical AI asks for your trust. This week, it had to show its work.
In one clinic, an algorithm faced the thing we keep saying every tool should face, a randomized controlled trial, and came out the other side with numbers that would make any diagnostician look twice. In another corner, doctors quietly out-prescribed a purpose-built medical model on the infections that actually keep you up at night. And in the background, the most popular "ChatGPT for doctors" spent the week fighting with a journal over whether its report card counts.
Proof, not promise. That's the theme this Sunday: what it looks like when the evidence finally arrives, what it costs when it doesn't, and who's doing the hard work of finding out.
Let's get into it.
- Troy, Ray, and Ibrahim

The AI finally showed its work

On July 24, Nature Medicine published something clinical AI has been promising for years and almost never delivers: a multicenter randomized controlled trial. The tool is Retina4IRD, a decision-support system for inherited retinal disease built on the Eye2Gene project at UCL and Moorfields Eye Hospital. Researchers randomized 300 patients with suspected inherited retinal disease to genetic workup by specialists working with the AI, or without it.
The result wasn't subtle. With AI support, specialists put the correct gene in their top five 88.5% of the time, versus 67.3% without. That's a 21-point jump. First-choice accuracy rose from 22.4% to 37.8%, and downstream management decisions scored significantly better too.
Why it matters: Nearly every diagnostic AI you've heard of was validated retrospectively, on benchmarks. This one changed what real specialists actually decided, prospectively and under randomization, in a disease where a genetic diagnosis is the gate to gene-therapy eligibility.
The catch: One disease domain, tertiary academic centers, and the AI assisted specialists rather than replacing anyone. Whether it holds up in a community clinic is untested. And the same week, a separate blinded evaluation found AI still trailing physicians on complex inpatient antibiotic decisions. Proof cuts both ways.
Bottom line: This is the bar. Not leaderboards, not demos. Trials that measure whether clinicians with the tool make better decisions than clinicians without it. More of this, please.
Honest check: would a randomized trial like this one change how much you trust AI at the bedside?

We're collecting stories. This week, one question.
This week's theme is proof, so let's collect some. Here's this week's question:
Tell us about the tool that actually earned its place at your bedside, or the one that arrived with fanfare and was quietly abandoned by spring. Maybe it caught something you'd have walked past. Maybe you're the nurse, resident, medic, or rollout lead who watched it quietly fail. Either way, you found out what the brochures never say.
We want the real version, not the conference-panel version. Two paragraphs is plenty.
This is open to every corner of medicine: EMS, nursing, pharmacy, techs, registration, environmental services, and administration, all of it. You choose how you're named, whether that's full name, role only, or fully anonymous. We protect patients in every story we run. That is not negotiable and it never will be.
If we run your story, we'll send you a Consult mug as a thank you.

Drs. Michel Michaelides & Nikolas Pontikos, UCL Institute of Ophthalmology & Moorfields Eye Hospital, London
Behind this week's feature is a clinician-scientist team that did it the slow way. Michel Michaelides is a professor of ophthalmology and a retinal specialist; Nikolas Pontikos leads Eye2Gene, which has been training algorithms on Moorfields' retinal imaging since 2017. Rather than shipping a benchmark and a press release, they took their system to the hardest test in medicine, randomization, and committed to publishing the result either way.
Where things stand:
Ran the first randomized controlled trial of AI-assisted genetic diagnosis in inherited retinal disease: 295 patients analyzed, multicenter, published in Nature Medicine this July
Top-5 diagnostic accuracy improved from 67.3% to 88.5% with their system in the loop
Built on Eye2Gene, backed by the UK Inherited Retinal Disease Consortium since 2017
Designed to assist specialists, not replace them, and it measurably improved management decisions, not just diagnoses

Illustration: The Consult
Why they matter: Nearly every AI headline this year came from a benchmark. This team volunteered for randomization instead, in a disease where the genetic answer decides who is eligible for gene therapy. If the field follows their lead, the next generation of clinical AI arrives with evidence instead of adjectives.


Neko Health lands $700M to scan America
Daniel Ek's Stockholm-based prevention startup closed a Series C at a roughly $7B valuation, about four times last year's, with its first US clinic planned for New York. The product: a 60-minute, non-invasive full-body scan reviewed with a physician. Screening's oldest question comes along for the ride: who does this actually help?

Medicare starts paying for AI like software
CMS is updating billing codes so clinical algorithms are reimbursed as software rather than as physical imaging equipment. It sounds bureaucratic; it isn't. How AI gets paid will quietly decide which tools survive contact with your health system's budget.

Five in six clinicians are winging it
A new industry roundup reports that roughly 5 in 6 clinicians using AI at work have no guidance from their employer on how to use it, and 29% of patient-AI conversations happen outside business hours. The tools arrived; the rulebook didn't.

The UK bets £75M on retiring the lab mouse
A government-backed program will test whether simulated-biology models can stand in for animal toxicology studies in drug-safety testing. If regulators accept the evidence, it would be one of the biggest changes to preclinical research in decades.

The evening, not the diagnosis
There's a line making the rounds this year: the most valuable thing AI can give a doctor in 2026 isn't a cleverer diagnosis. It's their evening back.
It's a good line, and this week shows why it's only half the story. The retina trial proved a machine can sharpen clinical judgment. The antibiotic study proved it still can't replace it. And the benchmark fight proved we barely agree on how to measure either. Meanwhile, the thing AI most reliably does in a hospital today is quieter: it writes the note, drafts the message, closes the chart before dinner instead of after.
That matters more than it sounds. Around 61% of healthcare workers report moderate-to-extreme burnout, and researchers increasingly split it in two: workload burnout (the documentation, the inbox, the pajama time) and what some call toxic burnout, the kind that comes from understaffing, unsafe conditions, and broken trust. Software can take real weight out of the first. It cannot touch the second.
Which is a clarifying way to read every announcement in this issue. The paperwork was never the job. It was the thing standing between you and the job: hearing the tremor in a voice, catching what the family isn't saying, owning the decision when the answer isn't in any dataset, human or machine.
If the machines take the charting and we get the evenings back, we're left with the oldest question in medicine: what do we do with the attention we've been given back? Whatever the benchmarks say, that answer still belongs to us.
Until next Sunday,

For the people keeping medicine human.
