The Opening

For two years, the loudest question in health AI was what the tools could do. This week, the question was who gets to set the rules.

In Washington, three federal offices quietly invited experts to a one-month sprint on how clinical AI should be graded. In the states, new laws came online telling insurers and chatbots what they may no longer do. And in a research journal, mock jurors weighed in on the question every clinician actually cares about: what happens to the doctor when the algorithm is wrong?

None of this makes headlines the way a flashy model launch does. All of it will decide what sits next to you at the bedside five years from now.

So this Sunday, we follow the rulebook: who is writing it, who is paying under it, and who answers for it when it fails.

Let's get into it.

- Troy, Ray, and Ibrahim

Washington wants a report card for clinical AI

Report card hovering over a stethoscope with a government dome in the background

On July 23, STAT reported that the White House Office of Science and Technology Policy, the FDA, and the HHS office that oversees health IT are convening outside experts to build something clinical AI has never had: shared standards for benchmarking and evaluating the tools clinicians actually use. The plan, laid out in an invitation reviewed by STAT's Mario Aguilar, is a one-month sprint, split into written and discussion phases, aimed at a consensus set of principles.

The timing is not an accident. The FDA's public list of AI-enabled devices has grown past 1,500 entries, and the agency's January guidance widened the category of decision-support tools it will not treat as devices at all. That leaves hospitals shopping for algorithms with no shared yardstick for whether they work.

Why it matters: Nearly every story in this newsletter eventually hits the same wall: nobody agrees on how to measure clinical AI. A common evaluation framework is the unglamorous infrastructure that turns marketing claims into answerable questions.

The catch: A month is a sprint, and consensus principles are not binding rules. Who gets a seat in the room, and whether model makers end up grading their own homework, will matter more than the announcement.

Bottom line: Regulation is arriving in pieces: guidance here, a state law there. A shared definition of good would be the piece that makes the rest of them work.

We're collecting stories. This week, one question.

This week's theme is the rulebook, so tell us about yours.

Tell us about the rule, the policy, or the total absence of one that shaped how AI showed up in your work. Maybe your hospital wrote a governance policy that actually helped. Maybe a tool appeared in your workflow one Tuesday with no memo, no training, and no one to ask. Either way, you lived the gap between the policy binder and the bedside.

We want the real version, not the conference-panel version. Two paragraphs is plenty.

This is open to every corner of medicine: EMS, nursing, pharmacy, techs, registration, environmental services, and administration, all of it. You choose how you're named, whether that's full name, role only, or fully anonymous. We protect patients in every story we run. That is not negotiable and it never will be.

If we run your story, we'll send you a Consult mug as a thank you.

Dr. Michelle Tarver, Director, FDA Center for Devices and Radiological Health

Every AI device that reaches an American bedside passes through the center Michelle Tarver runs. An ophthalmologist who still sees patients, she spent more than 15 years inside CDRH, built the FDA's first patient engagement advisory committee, and made a career of asking how devices land on actual people. Now she is the regulator deciding how algorithms that keep changing after clearance should be watched.

Where things stand:

  • Named permanent CDRH director in October 2024, after stepping in as acting director when her predecessor departed that July

  • Leads the center behind the FDA's list of AI-enabled medical devices, now past 1,500 entries, the overwhelming majority cleared through the 510(k) pathway

  • Told the AAMI neXus conference in April that final guidance on AI lifecycle management is coming, building on the draft published in January 2025

  • Has held two advisory committee meetings on generative AI and says to watch for the center's initial thinking by the end of 2026

Silhouetted figure at a podium addressing floating medical device icons

Illustration: The Consult

Why they matter: The feature above is about Washington looking for a yardstick. Tarver already holds the biggest one. Her stated bar is simple and hard: algorithms must be trained and validated on the populations they will actually serve, and watched after they ship. How her center finalizes that thinking will decide whether oversight of learning algorithms means real monitoring or paperwork.

State capitol with a gavel and circuit pattern between its columns

State AI rules just grew teeth

Colorado's first-in-the-nation law for high-risk AI took effect June 30, requiring impact assessments and notification when AI touches consequential clinical or financial decisions. Indiana followed on July 1, barring insurers from downcoding claims on AI's word alone. By one count, 21 states now have healthcare AI statutes on the books, which makes compliance, not capability, the new hard part.

Scale of justice balancing a computer chip and a white coat before silhouetted jurors

Jurors side with the algorithm

A randomized vignette trial in the Journal of Nuclear Medicine tested how mock jurors judge physicians who accept or reject AI advice. It flips the usual fear: jurors were more likely to find doctors liable when they rejected an AI recommendation for standard care. The catch is that vignettes are not verdicts, and no US law yet says who owns the error.

Magnifying glass over a chat bubble linked to an insulin pen and flowchart

The first "FDA-cleared AI agent" has fine print

UpDoc announced the first FDA-cleared clinical AI platform for real-time care delivery. A close read of the 510(k) by regulatory firm Innolitics found something narrower: an insulin-management tool where the language model only collects and relays information, while clinician-configured, deterministic logic makes the dosing calls. The gap between what AI is marketed as and what it is cleared to do is becoming its own story.

Globe orbited by a loop of connected nodes with a medical cross at the center

Global regulators sketch one AI rulebook

On July 10, the International Medical Device Regulators Forum, which includes the FDA and its counterpart agencies abroad, released a proposed technical framework for managing AI across its whole life cycle and opened it for public comment. Harmonized expectations would spare developers from reinventing compliance in every market. Consensus documents only bite, though, when national agencies actually adopt them.

The jury watches television too

This season of The Pitt gave its emergency department an AI problem. A new attending arrives with software promising to cut charting time by 80 percent. By episode six it is inventing patient details and confusing specialties, and a colleague delivers the line every hospital IT committee should laminate: "I need accurate information in the medical record." Later, a cyberattack knocks the systems offline, and the show quietly asks whether the residents still know how to work without them.

It is good drama. It is also, for millions of viewers, a curriculum. The people who become mock jurors in liability studies, who answer the polls legislators read, who will someday sit on an actual jury weighing an actual AI malpractice case, most of them will never open an FDA guidance document. They will have seen the show.

Medicine has been here before. Generations learned what a code looks like from television long before they saw a real one, and clinicians have spent careers gently correcting the expectations that came with that. Now television is drafting the public's assumptions about algorithms faster than any agency drafts guidance.

To their credit, the writers resisted the easy endings. The AI on The Pitt is neither savior nor villain; it is a tool that fails in ordinary ways while everyone is busy. That is roughly the truth, and roughly what the sprint in Washington is trying to write down as principles.

But when the real error comes, and it will, the jury's question will not be whether the algorithm met its benchmark. It will be the older one: was somebody paying attention to me? Every rulebook this week is really an argument about who holds that answer. So, who do we want holding it?

Until next Sunday,

For the people keeping medicine human.