The 30-second version
- When an AI handed radiologists the wrong answer, very experienced readers went from scoring 82.3% of mammograms correctly to 45.5%. The least experienced fell to 19.8%. Experience helped. It did not protect.
- And yet the best evidence is genuinely good: in a randomized trial of 105,934 women, AI-supported screening found more cancers with 44% less reading, with no rise in interval cancers.
- Lab performance does not transfer. The same class of tool that dazzles in a trial hit 35% sensitivity in a real primary-care population.
- So the answer is not "does AI work." It is does this model still work here, this month, which is a monitoring problem, and radiology has quietly started building the boring infrastructure to solve it.
The most important number I found while researching this series is not a sensitivity or an area under a curve. It is what happened to expert radiologists when the machine was confidently wrong.
This is the second of three posts. The first argued that the fear of AI eliminating jobs is aimed at the wrong target, because health care cannot staff the work it already has. That argument has a hole in it, and I want to put my finger in it before going any further: a staffing crisis is a reason to want a tool. It is not evidence that the tool works.
So this post is the evidence. All of it, including the parts I wish were different.
The short version is that the good trials are better than I expected and the failure modes are worse. Both of those things are true at once, and any version of this conversation that gives you only one of them is selling something.
What the good evidence actually shows
The strongest data we have comes from breast screening, because that is where somebody finally did the randomized trial.
The MASAI trial in Sweden randomized 105,934 women to either AI-supported screening or standard double reading by two radiologists. The AI triaged which exams needed a second reader and flagged suspicious findings. The final results, published in The Lancet in 2026, reported the primary outcome: the interval cancer rate, meaning cancers that surface between screening rounds because the screen missed them. That rate was 1.55 per 1,000 with AI versus 1.76 without, which met the trial's bar for non-inferiority. Sensitivity was higher with AI, 80.5% versus 73.8%. Specificity was identical at 98.5%.
The workload number is the one I keep coming back to. AI-supported screening required 61,248 screen readings where standard double reading required 109,692. That is a 44% reduction in reading, with more cancers found and no significant increase in false positives.
Germany replicated the direction at enormous scale. The PRAIM study followed 463,094 women screened by 119 radiologists across 12 sites and found a cancer detection rate of 6.7 per 1,000 with AI support versus 5.7 without, a 17.6% relative increase, with a recall rate that was slightly lower rather than higher. One important caveat: PRAIM was observational and the radiologists chose for themselves whether to use the AI, so the groups were not randomly assigned and the comparison is weaker than MASAI's.
Notice what actually improved. Not the radiologist's eye. The radiologist's throughput, and the number of second reads that never needed a human at all.
The place this matters most is not Scottsdale
Everything above happened in wealthy countries with organized screening programs and plenty of radiologists. The larger prize is somewhere else entirely.
In Bangladesh, researchers ran 23,954 chest X-rays from three tuberculosis screening centers past five commercial AI algorithms and a panel of three registered radiologists. All five algorithms significantly outperformed the radiologists, and all five cut the number of confirmatory molecular tests needed by about half while holding sensitivity above 90%. The authors also reported that every algorithm performed worse in people over 60 and in people with a history of TB, which is exactly the kind of subgroup detail that gets dropped when these results are summarized.
In China, a study across 7 county-level and 32 township-level facilities reviewed 93,319 patients, of whom 273 had bacteriologically confirmed pulmonary TB. The AI flagged 83.9% of those confirmed cases; the radiologists reading at the time caught 25.6%. The AI's positive predictive value was much worse, 1.7% against 10.3%, meaning far more false alarms. But used as a triage filter with human review of flagged images, it cut the radiologist workload by 85.5% without missing any case the radiologists had found on their own.
A separate validation on more than one million chest X-rays reported an AUC of 98.51% and a false negative rate slightly better than the radiologists', with the potential to auto-report up to 80% of normal studies.
These are the numbers that make me hopeful, and they have almost nothing to do with whether AI is better than me. They are about places where there is no radiologist to be better than. That is where "raising all boats" stops being a slogan.
Four findings that should keep us honest
I do not want to write a brochure. Here is the evidence that cuts the other way, and some of it is genuinely alarming.
Automation bias is worse than I expected. In a prospective experiment, 27 radiologists read mammograms with a purported AI system that was deliberately wrong on 12 of 40 cases. Among the most experienced readers, the share of correctly assigned BI-RADS categories fell from 82.3% to 45.5% when the AI suggested the wrong category. Among the least experienced, it fell from 79.7% to 19.8%. Experience helped. It did not protect.
Deskilling may be real. Four Polish endoscopy centers compared adenoma detection during unassisted colonoscopy in the three months before AI was introduced and the three months after. The rate fell from 28.4% to 22.4%, an absolute drop of 6 percentage points. This was a retrospective before-and-after comparison, not a randomized one, so seasonality, case mix, and staffing changes are all live alternative explanations. But the effect size is large enough that dismissing it would be motivated reasoning.
Lab performance does not transfer automatically. A commercial chest X-ray AI validated against CT findings in 3,047 consecutive radiographs from two primary healthcare centers, where the prevalence of significant disease was 2.2%, achieved a sensitivity of 35.3% and an AUROC of 0.648. The authors' conclusion is the sentence I would put on a poster in every department: regulatory approval and experimental performance may not translate to real practice, and the mismatch tends to be worst exactly where the need is greatest.
AI did not make radiologists less burned out. A survey of 6,726 radiologists across 1,143 Chinese hospitals found that frequent AI users had higher odds of burnout than non-users, with an adjusted odds ratio of 1.20 and a dose-response relationship with frequency of use, driven mostly by emotional exhaustion. It was worst among radiologists with high workloads. This is cross-sectional, so causality could run either way. But it points at something that matches my experience: if you speed up one task and leave the volume expectation untouched, you have not reduced anyone's suffering. You have just changed what they do all day and asked for more of it.
Equity does not happen by itself
The version of this future I want is one where a woman in a rural county gets her MRI read this week instead of in March. But the technology does not deliver that on its own, and there is a well-documented case showing exactly how it fails.
A commercial algorithm used across US health systems to identify patients needing extra care was found to be substantially biased against Black patients: at any given risk score, Black patients were considerably sicker than white patients. The mechanism was not malice or a bad training set in the usual sense. The algorithm predicted health care costs as a proxy for illness, and because less money has historically been spent on Black patients, the proxy encoded the disparity. Correcting it would have raised the share of Black patients flagged for additional help from 17.7% to 46.5%.
That is a design decision, not an accident of the math. Somebody chose a convenient outcome variable. The same choice is available to every group building an imaging model right now, and it will be made well or badly depending on who is in the room.
My own specialty has actually built something
Here is the part of this story I did not expect to be writing, and the part I am proudest of. While the broader AI conversation has been arguing about whether guardrails are even possible, radiology quietly went and built some.
In June 2024 the American College of Radiology launched ARCH-AI, the ACR Recognized Center for Healthcare-AI, described as the first national quality assurance program for AI in medical imaging. To earn the designation, a practice attests to a specific set of things: that it has an interdisciplinary AI governance group, that it keeps a documented inventory of every algorithm it runs, that it has a deliberate process for reviewing and selecting those algorithms, that it does acceptance testing before deployment, that it monitors performance afterward, and that it manages any models it built itself.
None of that is glamorous. All of it is exactly what was missing in the failure modes above.
ARCH-AI is deliberately a stepping stone. The ACR leadership behind it have written openly that it exists as a precursor to a formal accreditation program, on the same model the College has used since radiation oncology in 1966 and mammography in 1987, with council approval anticipated around spring 2027. Their stated reason for building it is the same observation this whole post keeps circling: real-world AI performance often differs from what premarket testing showed. That sentence is in the ACR's own road map paper. It is not a criticism from outside the field.
The piece I find genuinely impressive is the second one. In November 2024 the ACR launched Assess-AI, a registry inside the National Radiology Data Registry that monitors how deployed imaging AI is actually performing, in real practices, over time. Participating sites send de-identified algorithm outputs, report text, and study metadata. The registry extracts surrogate labels from the radiology reports, computes concordance between what the algorithm said and what the radiologist ultimately said, and returns it as dashboards. Sites can compare themselves against national benchmarks and against peers matched on facility type, region, trauma level, and urban versus rural. They can drill into discordant cases and look at whether the disagreements cluster by demographic or technical factor. It currently covers intracranial hemorrhage, pulmonary embolism, pneumothorax, large-vessel occlusion, bone age, and cervical spine fracture.
Sit with what that is for a moment. It is post-market surveillance for algorithms, built by the specialty that uses them, that measures whether a model still works in your department, on your scanners, with your patients, after the vendor demo is over. Model drift is not hypothetical; departments change their protocols, equipment, and case mix constantly, and performance moves with them. Assess-AI is the mechanism for noticing.
In 2026 the ACR and SIIM also approved a formal practice parameter covering tool selection, predeployment evaluation, ongoing monitoring, and privacy, and the program has begun expanding internationally, with the University Hospital of Bern named its first site outside the United States.
Radiology did not wait to be regulated. It built the registry, wrote the parameter, and put a badge on the wall. I would like the rest of the AI industry to notice that this was possible.
I want to be measured about it. ARCH-AI is attestation, not audit: a practice affirms it is doing these things rather than being inspected. Assess-AI depends on voluntary participation and on surrogate labels pulled from report text by a language model, which is a reasonable approximation of truth and not truth itself. Neither program stops a department from buying a bad algorithm. What they do is make it much harder to buy one and never find out.
Before your department buys an imaging AI, ask these
ARCH-AI covers the institutional layer. These are the clinical questions underneath it.
- What population was it validated on, and what was the disease prevalence in that population compared to ours?
- What are the reported subgroup results by age, sex, race, body habitus, and scanner vendor? If there are none, that is the answer.
- Is it a triage tool, a second reader, or a concurrent reader? Each one fails differently and each one needs a different workflow.
- What happens to the reading list when it is wrong, and how would we ever find out that it was?
- Does the volume expectation change when the tool goes live? If throughput goes up and staffing does not, we have bought a burnout accelerator.
- Are we submitting to Assess-AI, and if not, what is our alternative plan for catching drift?
- How do we preserve the skills of residents and junior attendings who will train alongside it?
What this post does not tell you
Two posts in, I have argued that the work is moving rather than vanishing, and that the technology is real but conditional: good in the trial, fragile in the field, safe only with a loop and a registry behind it.
Both of those are arguments about what medicine should do. They assume medicine gets to decide the pace.
I no longer think that is true. Several hundred million people a week are already asking these systems health questions, and a growing number have connected their own medical records to them. The knowledge asymmetry that defined the exam room for a century is closing from the patient's side, and nobody asked us.
That is the last post.
How this piece was built
Every trial number here was checked against the source abstract rather than a summary of it, and where a study is observational, small, or before-and-after rather than randomized, I have said so in the same sentence as the finding rather than in a footnote. Two figures in this post carry warnings about how to read them; please read them.
AI disclosure. The argument and the point of view are mine. I worked with Claude (Anthropic) as a research and drafting partner: it searched PubMed and Consensus, verified each trial's numbers against the primary source, produced the two figures, and helped organize the draft. I reviewed every claim and citation before publishing.
References
- Hernström V, et al. Screening performance and characteristics of breast cancer detected in the Mammography Screening with Artificial Intelligence trial (MASAI): a randomised, controlled, parallel-group, non-inferiority, single-blinded, screening accuracy study. Lancet Digit Health. 2025;7(3):e175-e183. doi:10.1016/S2589-7500(24)00267-X
- Hernström V, et al. Interval cancers and screening outcomes in the MASAI trial. Lancet. 2026;407(10430):505-514. doi:10.1016/S0140-6736(25)02464-X
- Lång K, et al. Artificial intelligence-supported screen reading versus standard double reading in the Mammography Screening with Artificial Intelligence trial (MASAI): a clinical safety analysis. Lancet Oncol. 2023;24(8):936-944. doi:10.1016/S1470-2045(23)00298-X
- Eisemann N, et al. Nationwide real-world implementation of AI for cancer detection in population-based mammography screening. Nat Med. 2025;31(3):917-924. doi:10.1038/s41591-024-03408-6
- Dratsch T, et al. Automation bias in mammography: the impact of artificial intelligence BI-RADS suggestions on reader performance. Radiology. 2023;307(4):e222176. doi:10.1148/radiol.222176
- Budzyń K, et al. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. Lancet Gastroenterol Hepatol. 2025;10(10):896-903. doi:10.1016/S2468-1253(25)00133-5
- Kim JH, et al. Clinical validation of a deep learning-based software for lung nodule detection in chest radiographs in a health screening population. Eur Radiol. 2023;33(11):7823-7833. doi:10.1007/s00330-023-09761-3
- Liu Y, et al. Artificial intelligence use and burnout among radiologists in China. JAMA Netw Open. 2024;7(12):e2448714. doi:10.1001/jamanetworkopen.2024.48714
- Qin ZZ, et al. Tuberculosis detection from chest x-rays for triaging in a high tuberculosis-burden setting: an evaluation of five artificial intelligence algorithms. Lancet Digit Health. 2021;3(9):e543-e554. doi:10.1016/S2589-7500(21)00116-3
- Jiang Y, et al. Effectiveness of computer-aided detection for active pulmonary tuberculosis screening in resource-limited settings. J Med Internet Res. 2025;27:e69109. doi:10.2196/69109
- Munjal P, et al. Assessing the reliability of an AI-based chest radiograph interpretation system on over one million radiographs. NPJ Digit Med. 2025;8(1):318. doi:10.1038/s41746-025-01693-0
- Obermeyer Z, et al. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366(6464):447-453. doi:10.1126/science.aax2342
- Larson DB, et al. A road map for the ACR Recognized Center for Healthcare-AI. J Am Coll Radiol. 2025;22(5):586-592. doi:10.1016/j.jacr.2025.02.008
- Coombs LP, et al. The ACR Assess-AI registry: national performance monitoring of clinical AI. J Am Coll Radiol. 2026;23(9):1557-1566. doi:10.1016/j.jacr.2026.04.024
Peer-reviewed sources were located through PubMed and Consensus. Nothing in this post describes any individual patient or protected health information, and nothing here represents the position of my employer.
No comments:
Post a Comment