Tuesday, September 15, 2026

The Good Trials Are Better Than I Expected. The Failure Modes Are Worse.

The 30-second version

  • When an AI handed radiologists the wrong answer, very experienced readers went from scoring 82.3% of mammograms correctly to 45.5%. The least experienced fell to 19.8%. Experience helped. It did not protect.
  • And yet the best evidence is genuinely good: in a randomized trial of 105,934 women, AI-supported screening found more cancers with 44% less reading, with no rise in interval cancers.
  • Lab performance does not transfer. The same class of tool that dazzles in a trial hit 35% sensitivity in a real primary-care population.
  • So the answer is not "does AI work." It is does this model still work here, this month, which is a monitoring problem, and radiology has quietly started building the boring infrastructure to solve it.

The most important number I found while researching this series is not a sensitivity or an area under a curve. It is what happened to expert radiologists when the machine was confidently wrong.

This is the second of three posts. The first argued that the fear of AI eliminating jobs is aimed at the wrong target, because health care cannot staff the work it already has. That argument has a hole in it, and I want to put my finger in it before going any further: a staffing crisis is a reason to want a tool. It is not evidence that the tool works.

So this post is the evidence. All of it, including the parts I wish were different.

The short version is that the good trials are better than I expected and the failure modes are worse. Both of those things are true at once, and any version of this conversation that gives you only one of them is selling something.

What the good evidence actually shows

The strongest data we have comes from breast screening, because that is where somebody finally did the randomized trial.

The MASAI trial in Sweden randomized 105,934 women to either AI-supported screening or standard double reading by two radiologists. The AI triaged which exams needed a second reader and flagged suspicious findings. The final results, published in The Lancet in 2026, reported the primary outcome: the interval cancer rate, meaning cancers that surface between screening rounds because the screen missed them. That rate was 1.55 per 1,000 with AI versus 1.76 without, which met the trial's bar for non-inferiority. Sensitivity was higher with AI, 80.5% versus 73.8%. Specificity was identical at 98.5%.

Two radiologists, no AI AI-supported reading
Cancers found per 1,000 women screened 5.0 6.4 Sensitivity 73.8% 80.5% Screen readings the radiologists had to do 109,692 61,248 Compare bars within a panel, never across panels. Bar lengths are scaled per panel.
One randomized trial, 105,934 women, Sweden. More cancers found, higher sensitivity, same specificity, and 44% less reading. The three panels use different scales, so only the within-panel comparison is meaningful. Detection and workload figures from the 2025 Lancet Digital Health report; sensitivity from the 2026 Lancet primary analysis of the same trial.

The workload number is the one I keep coming back to. AI-supported screening required 61,248 screen readings where standard double reading required 109,692. That is a 44% reduction in reading, with more cancers found and no significant increase in false positives.

Germany replicated the direction at enormous scale. The PRAIM study followed 463,094 women screened by 119 radiologists across 12 sites and found a cancer detection rate of 6.7 per 1,000 with AI support versus 5.7 without, a 17.6% relative increase, with a recall rate that was slightly lower rather than higher. One important caveat: PRAIM was observational and the radiologists chose for themselves whether to use the AI, so the groups were not randomly assigned and the comparison is weaker than MASAI's.

Notice what actually improved. Not the radiologist's eye. The radiologist's throughput, and the number of second reads that never needed a human at all.

The place this matters most is not Scottsdale

Everything above happened in wealthy countries with organized screening programs and plenty of radiologists. The larger prize is somewhere else entirely.

In Bangladesh, researchers ran 23,954 chest X-rays from three tuberculosis screening centers past five commercial AI algorithms and a panel of three registered radiologists. All five algorithms significantly outperformed the radiologists, and all five cut the number of confirmatory molecular tests needed by about half while holding sensitivity above 90%. The authors also reported that every algorithm performed worse in people over 60 and in people with a history of TB, which is exactly the kind of subgroup detail that gets dropped when these results are summarized.

In China, a study across 7 county-level and 32 township-level facilities reviewed 93,319 patients, of whom 273 had bacteriologically confirmed pulmonary TB. The AI flagged 83.9% of those confirmed cases; the radiologists reading at the time caught 25.6%. The AI's positive predictive value was much worse, 1.7% against 10.3%, meaning far more false alarms. But used as a triage filter with human review of flagged images, it cut the radiologist workload by 85.5% without missing any case the radiologists had found on their own.

A separate validation on more than one million chest X-rays reported an AUC of 98.51% and a false negative rate slightly better than the radiologists', with the potential to auto-report up to 80% of normal studies.

These are the numbers that make me hopeful, and they have almost nothing to do with whether AI is better than me. They are about places where there is no radiologist to be better than. That is where "raising all boats" stops being a slogan.

Four findings that should keep us honest

I do not want to write a brochure. Here is the evidence that cuts the other way, and some of it is genuinely alarming.

OR 1.20
Higher odds of burnout among radiologists who used AI frequently, with a dose-response by frequency of use
6,726 radiologists, 1,143 hospitals (Liu 2024)
−6.0pts
Drop in adenoma detection when endoscopists went back to working without AI, after months of using it
1,443 colonoscopies, 4 centers (Budzyń 2025)
35%
Sensitivity of a commercial chest X-ray AI in a real-world, low-prevalence screening population
3,047 radiographs, 2 primary care centers (Kim 2023)

Automation bias is worse than I expected. In a prospective experiment, 27 radiologists read mammograms with a purported AI system that was deliberately wrong on 12 of 40 cases. Among the most experienced readers, the share of correctly assigned BI-RADS categories fell from 82.3% to 45.5% when the AI suggested the wrong category. Among the least experienced, it fell from 79.7% to 19.8%. Experience helped. It did not protect.

AI suggested the correct category AI suggested the wrong one
Mammograms assigned the correct BI-RADS category 27 radiologists, 50 cases, AI deliberately wrong on some of them Very experienced 82.3% 45.5% Moderately experienced 81.3% 24.8% Inexperienced 79.7% 19.8% 0% 100% Read the blue bars first: all three groups start in the same place. Then read the gold ones.
Experience buys you some protection. Not much. All six bars are zero-based on one shared scale, so every length is directly comparable. The three groups perform almost identically when the machine is right. The gap opens only when it is wrong. This was a prospective experiment with a purported AI system, not a deployed product, which makes it a clean measure of the effect and not a claim about any commercial tool. Dratsch et al., Radiology 2023.

Deskilling may be real. Four Polish endoscopy centers compared adenoma detection during unassisted colonoscopy in the three months before AI was introduced and the three months after. The rate fell from 28.4% to 22.4%, an absolute drop of 6 percentage points. This was a retrospective before-and-after comparison, not a randomized one, so seasonality, case mix, and staffing changes are all live alternative explanations. But the effect size is large enough that dismissing it would be motivated reasoning.

Lab performance does not transfer automatically. A commercial chest X-ray AI validated against CT findings in 3,047 consecutive radiographs from two primary healthcare centers, where the prevalence of significant disease was 2.2%, achieved a sensitivity of 35.3% and an AUROC of 0.648. The authors' conclusion is the sentence I would put on a poster in every department: regulatory approval and experimental performance may not translate to real practice, and the mismatch tends to be worst exactly where the need is greatest.

AI did not make radiologists less burned out. A survey of 6,726 radiologists across 1,143 Chinese hospitals found that frequent AI users had higher odds of burnout than non-users, with an adjusted odds ratio of 1.20 and a dose-response relationship with frequency of use, driven mostly by emotional exhaustion. It was worst among radiologists with high workloads. This is cross-sectional, so causality could run either way. But it points at something that matches my experience: if you speed up one task and leave the volume expectation untouched, you have not reduced anyone's suffering. You have just changed what they do all day and asked for more of it.

Equity does not happen by itself

The version of this future I want is one where a woman in a rural county gets her MRI read this week instead of in March. But the technology does not deliver that on its own, and there is a well-documented case showing exactly how it fails.

A commercial algorithm used across US health systems to identify patients needing extra care was found to be substantially biased against Black patients: at any given risk score, Black patients were considerably sicker than white patients. The mechanism was not malice or a bad training set in the usual sense. The algorithm predicted health care costs as a proxy for illness, and because less money has historically been spent on Black patients, the proxy encoded the disparity. Correcting it would have raised the share of Black patients flagged for additional help from 17.7% to 46.5%.

That is a design decision, not an accident of the math. Somebody chose a convenient outcome variable. The same choice is available to every group building an imaging model right now, and it will be made well or badly depending on who is in the room.

My own specialty has actually built something

Here is the part of this story I did not expect to be writing, and the part I am proudest of. While the broader AI conversation has been arguing about whether guardrails are even possible, radiology quietly went and built some.

In June 2024 the American College of Radiology launched ARCH-AI, the ACR Recognized Center for Healthcare-AI, described as the first national quality assurance program for AI in medical imaging. To earn the designation, a practice attests to a specific set of things: that it has an interdisciplinary AI governance group, that it keeps a documented inventory of every algorithm it runs, that it has a deliberate process for reviewing and selecting those algorithms, that it does acceptance testing before deployment, that it monitors performance afterward, and that it manages any models it built itself.

None of that is glamorous. All of it is exactly what was missing in the failure modes above.

ARCH-AI is deliberately a stepping stone. The ACR leadership behind it have written openly that it exists as a precursor to a formal accreditation program, on the same model the College has used since radiation oncology in 1966 and mammography in 1987, with council approval anticipated around spring 2027. Their stated reason for building it is the same observation this whole post keeps circling: real-world AI performance often differs from what premarket testing showed. That sentence is in the ACR's own road map paper. It is not a criticism from outside the field.

The piece I find genuinely impressive is the second one. In November 2024 the ACR launched Assess-AI, a registry inside the National Radiology Data Registry that monitors how deployed imaging AI is actually performing, in real practices, over time. Participating sites send de-identified algorithm outputs, report text, and study metadata. The registry extracts surrogate labels from the radiology reports, computes concordance between what the algorithm said and what the radiologist ultimately said, and returns it as dashboards. Sites can compare themselves against national benchmarks and against peers matched on facility type, region, trauma level, and urban versus rural. They can drill into discordant cases and look at whether the disagreements cluster by demographic or technical factor. It currently covers intracranial hemorrhage, pulmonary embolism, pneumothorax, large-vessel occlusion, bone age, and cervical spine fracture.

Sit with what that is for a moment. It is post-market surveillance for algorithms, built by the specialty that uses them, that measures whether a model still works in your department, on your scanners, with your patients, after the vendor demo is over. Model drift is not hypothetical; departments change their protocols, equipment, and case mix constantly, and performance moves with them. Assess-AI is the mechanism for noticing.

In 2026 the ACR and SIIM also approved a formal practice parameter covering tool selection, predeployment evaluation, ongoing monitoring, and privacy, and the program has begun expanding internationally, with the University Hospital of Bern named its first site outside the United States.

Radiology did not wait to be regulated. It built the registry, wrote the parameter, and put a badge on the wall. I would like the rest of the AI industry to notice that this was possible.

I want to be measured about it. ARCH-AI is attestation, not audit: a practice affirms it is doing these things rather than being inspected. Assess-AI depends on voluntary participation and on surrogate labels pulled from report text by a language model, which is a reasonable approximation of truth and not truth itself. Neither program stops a department from buying a bad algorithm. What they do is make it much harder to buy one and never find out.

Before your department buys an imaging AI, ask these

ARCH-AI covers the institutional layer. These are the clinical questions underneath it.

  • What population was it validated on, and what was the disease prevalence in that population compared to ours?
  • What are the reported subgroup results by age, sex, race, body habitus, and scanner vendor? If there are none, that is the answer.
  • Is it a triage tool, a second reader, or a concurrent reader? Each one fails differently and each one needs a different workflow.
  • What happens to the reading list when it is wrong, and how would we ever find out that it was?
  • Does the volume expectation change when the tool goes live? If throughput goes up and staffing does not, we have bought a burnout accelerator.
  • Are we submitting to Assess-AI, and if not, what is our alternative plan for catching drift?
  • How do we preserve the skills of residents and junior attendings who will train alongside it?

What this post does not tell you

Two posts in, I have argued that the work is moving rather than vanishing, and that the technology is real but conditional: good in the trial, fragile in the field, safe only with a loop and a registry behind it.

Both of those are arguments about what medicine should do. They assume medicine gets to decide the pace.

I no longer think that is true. Several hundred million people a week are already asking these systems health questions, and a growing number have connected their own medical records to them. The knowledge asymmetry that defined the exam room for a century is closing from the patient's side, and nobody asked us.

That is the last post.


How this piece was built

Every trial number here was checked against the source abstract rather than a summary of it, and where a study is observational, small, or before-and-after rather than randomized, I have said so in the same sentence as the finding rather than in a footnote. Two figures in this post carry warnings about how to read them; please read them.

AI disclosure. The argument and the point of view are mine. I worked with Claude (Anthropic) as a research and drafting partner: it searched PubMed and Consensus, verified each trial's numbers against the primary source, produced the two figures, and helped organize the draft. I reviewed every claim and citation before publishing.

References

  1. Hernström V, et al. Screening performance and characteristics of breast cancer detected in the Mammography Screening with Artificial Intelligence trial (MASAI): a randomised, controlled, parallel-group, non-inferiority, single-blinded, screening accuracy study. Lancet Digit Health. 2025;7(3):e175-e183. doi:10.1016/S2589-7500(24)00267-X
  2. Hernström V, et al. Interval cancers and screening outcomes in the MASAI trial. Lancet. 2026;407(10430):505-514. doi:10.1016/S0140-6736(25)02464-X
  3. Lång K, et al. Artificial intelligence-supported screen reading versus standard double reading in the Mammography Screening with Artificial Intelligence trial (MASAI): a clinical safety analysis. Lancet Oncol. 2023;24(8):936-944. doi:10.1016/S1470-2045(23)00298-X
  4. Eisemann N, et al. Nationwide real-world implementation of AI for cancer detection in population-based mammography screening. Nat Med. 2025;31(3):917-924. doi:10.1038/s41591-024-03408-6
  5. Dratsch T, et al. Automation bias in mammography: the impact of artificial intelligence BI-RADS suggestions on reader performance. Radiology. 2023;307(4):e222176. doi:10.1148/radiol.222176
  6. Budzyń K, et al. Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy: a multicentre, observational study. Lancet Gastroenterol Hepatol. 2025;10(10):896-903. doi:10.1016/S2468-1253(25)00133-5
  7. Kim JH, et al. Clinical validation of a deep learning-based software for lung nodule detection in chest radiographs in a health screening population. Eur Radiol. 2023;33(11):7823-7833. doi:10.1007/s00330-023-09761-3
  8. Liu Y, et al. Artificial intelligence use and burnout among radiologists in China. JAMA Netw Open. 2024;7(12):e2448714. doi:10.1001/jamanetworkopen.2024.48714
  9. Qin ZZ, et al. Tuberculosis detection from chest x-rays for triaging in a high tuberculosis-burden setting: an evaluation of five artificial intelligence algorithms. Lancet Digit Health. 2021;3(9):e543-e554. doi:10.1016/S2589-7500(21)00116-3
  10. Jiang Y, et al. Effectiveness of computer-aided detection for active pulmonary tuberculosis screening in resource-limited settings. J Med Internet Res. 2025;27:e69109. doi:10.2196/69109
  11. Munjal P, et al. Assessing the reliability of an AI-based chest radiograph interpretation system on over one million radiographs. NPJ Digit Med. 2025;8(1):318. doi:10.1038/s41746-025-01693-0
  12. Obermeyer Z, et al. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366(6464):447-453. doi:10.1126/science.aax2342
  13. Larson DB, et al. A road map for the ACR Recognized Center for Healthcare-AI. J Am Coll Radiol. 2025;22(5):586-592. doi:10.1016/j.jacr.2025.02.008
  14. Coombs LP, et al. The ACR Assess-AI registry: national performance monitoring of clinical AI. J Am Coll Radiol. 2026;23(9):1557-1566. doi:10.1016/j.jacr.2026.04.024

Peer-reviewed sources were located through PubMed and Consensus. Nothing in this post describes any individual patient or protected health information, and nothing here represents the position of my employer.

Wednesday, September 9, 2026

Ten Years Ago, They Told Us to Stop Training Radiologists

The 30-second version

  • In 2016 a Nobel laureate told the world to stop training radiologists. There are more radiologists now than there were then, and the wait for an MRI is still measured in months.
  • Everyone tells the ATM story half-finished. Teller jobs did rise, and are now projected to fall 13.2%. Cashiers are down roughly 700,000 since 2019 and still falling.
  • But radiology is not retail, and the reason is demand, not prestige. Retail demand is roughly fixed, so faster checkout means fewer checkers. Imaging demand is nowhere close to met. Free capacity in a system with a backlog gets absorbed, not shed.
  • In one BLS projection round, office administration sheds 752,100 jobs while health care adds 2.2 million, about 37% of all new jobs in the economy. The work is moving, not vanishing. The fear is aimed at the wrong target.

Ten years ago a Nobel laureate told the world to stop training radiologists. I am a radiologist. I am still here, and so is everyone I trained with, and the waiting list for an MRI in this country is still months long. That gap between the prediction and the reality is worth understanding, because we are about to make the same category of mistake again.

In 2016, Geoffrey Hinton stood at a seminar in Toronto and said that anyone working as a radiologist was "like the coyote that's already over the edge of the cliff" and had not yet looked down. He said people should stop training radiologists immediately, because deep learning would beat them within five years. He is one of the most important scientists of the century and he was not being cynical. He genuinely believed it.

Medical students believed it too. Some of them changed their plans.

Here is where things actually landed. The number of practicing radiologists in the United States rose 12% between 2010 and 2022, from 34,328 to 38,306, which works out to a move from 11.1 to 11.5 radiologists per 100,000 people. One large midwestern academic medical center has reportedly grown its radiology staff by 55% since 2016, to roughly 400 radiologists, though that figure comes from press reporting rather than a peer-reviewed source. Meanwhile, imaging volume kept climbing and turnaround times kept stretching. The prediction did not just miss. It missed in the opposite direction.

The interesting part is who is quoting the story now. Jensen Huang has spent the last year telling audiences at Davos and elsewhere that radiology is the proof that AI creates jobs rather than destroying them. I want to be careful here, because a lot of people have heard "Jensen Huang says AI is coming for radiology" and that is backwards; he is arguing the opposite. But his version has its own problem. He claims hospitals are hiring more radiologists because of AI. That causal story is not supported. Imaging volumes have been growing steadily for decades for reasons that have nothing to do with machine learning, and interpretation speed has barely moved. We are hiring because we are drowning, not because algorithms freed us up.

Both the doom version and the triumph version get radiology wrong in the same way. They treat the job as a task, and the task as the job.

What I think, and where this series is going

This is the first of three posts, and I want to state my position rather than meander toward it. I think this is the most opportune moment in my professional lifetime to change how care is actually delivered, and that almost nobody is arguing about the right thing. The public fight is over whether a model can out-read a human. What will actually decide whether patients are better off is whether we can finally match expertise to need.

Across the three posts, here is what I believe:

  • The work redistributes rather than disappears. Health care cannot fill the jobs it already has. The fear of mass unemployment and the reality of a staffing crisis are unfolding in the same room, and the thing standing between them is retraining nobody has funded. That is this post.
  • Triage becomes the organizing principle, and the pair outperforms either half, but only if the human is genuinely in the loop and somebody is monitoring the machine forever. That is the second post, on what the evidence actually shows.
  • Reach matters more than accuracy, and patients level up the most. The knowledge asymmetry that defined the exam room for a century is closing from the patient's side, whether or not medicine approves. That is the third.

None of it is inevitable. It is a choice about what we build, not a forecast we are waiting on.

The ATM story, and the half of it nobody tells

Whenever this argument comes up, someone brings up bank tellers. It is the most-cited analogy in the automation debate and it is worth getting right.

The economist James Bessen went and collected the data. As ATMs spread across the United States, the number of tellers needed to run an average urban branch fell from about 20 to 13 between 1988 and 2004. But a cheaper branch is a branch worth opening, so banks opened 43% more of them in urban areas. Total teller employment held steady and even rose. The remaining work shifted toward relationships and problem-solving, the parts a cash dispenser could not touch.

That is the version everyone tells. Here is the part that gets left out: it did not last. The Bureau of Labor Statistics now projects teller employment to keep falling, because online and mobile banking eventually automated not just the cash-handling task but most of the reasons to walk into a branch at all.

Cashiers are the sharper case, and it is worth being concrete, because people wave at self-checkout without ever citing a figure. In May 2019 there were roughly 3.60 million cashiers in the United States, one of the largest occupations here. By 2025 there were 3.11 million. The BLS projects 2.91 million by 2035, a further loss of 200,600 jobs and a 6.5% decline, still the largest projected drop of any occupation in the country. Tellers are on that same table now, projected to fall 13.2%.

That is close to 700,000 cashier jobs gone or going in fifteen years. Not zero. Not a myth. The largest.

So the two stories are not the same story, and the difference is the whole point. Tellers rose while machines handled part of the job and are only now sliding, decades later, once phones removed the reason to enter a branch at all. Cashiers are further down that same road: the scanner and the payment terminal took the core of the work, the customer absorbed the rest, and one attendant now watches six lanes where four people used to stand.

The real lesson is not "automation never displaces anyone." It is a lesson about how much of a job gets automated, and it has a shape.

Machines take over some of the tasks Machines take over nearly all of them People employed in the occupation occupations travel left to right over time Bank tellers, 1990s ATMs cut staff per branch from 20 to 13, so banks opened 43% more branches. Total teller jobs went up. Radiology, today? Reading is a task. Diagnosis is the job. Bank tellers, now Phones removed the reason to enter a branch. Now falling. Cashiers, now Further along the same road, and falling fastest of all. What actually happened to cashiers 3.60M 3.11M 2.91M 2019 2025 2035, projected
The question is not whether automation displaces people. It is where on the curve you are standing, and which way you are moving. The curve itself is a conceptual diagram drawn to organize the argument; nobody has measured this shape directly, and the horizontal axis has no units. The cashier figures below it are real: about 3.60 million in May 2019 (BLS Occupational Employment and Wage Statistics), 3.11 million in 2025 and a projected 2.91 million in 2035 (BLS National Employment Matrix, 2025–2035 round). Those come from two different BLS programs with slightly different methods, so read the trend rather than the exact differences. Teller figures from Bessen's IMF analysis.

So which part of the curve is radiology on? Before I answer that, there is a second half to the cashier story that the employment figures do not capture, and I notice it every single week.

I used to dread Costco. The checkout line was the thing you planned the trip around, the reason you talked yourself out of going. Now I walk through it. The scanners, the app, the reconfigured front end: the experience is dramatically better, for me and honestly for the person working there, who is solving problems instead of dragging four hundred items across a piece of glass. The company benefits, the customer benefits, and the work that remains is more interesting than the work that went.

Both things are true at once. The experience got much better and several hundred thousand of those jobs went away. I am not going to pretend otherwise to make my argument tidier.

But here is the structural difference, and it is the hinge of everything that follows. Retail demand is roughly fixed. There is a certain amount of shopping to be done in a week, so making checkout twice as fast means needing about half as many checkers. Imaging demand is not fixed, and it is nowhere close to met. There are scans sitting unread right now. There are people waiting four months for an MRI. There are patients whose scan will never be ordered at all because the queue makes ordering it pointless, and whose disease will therefore be found later than it should have been.

When you free capacity in a system with a fixed amount of work, you shed people. When you free capacity in a system drowning in unserved need, the capacity gets absorbed. Radiology is the second kind, and that is an argument about demand, not about how special radiologists are.

Which means the question that matters is not where we sit on that curve. It is how big the backlog is. So let me show you.

The constraint is not accuracy. It is capacity.

Almost every public argument about AI and radiology is an argument about accuracy: can the model see the nodule, can it beat the human. That is the wrong axis. Accuracy is not what is failing patients right now.

What is failing patients is that there is not enough of us, and there is more and more imaging.

A study of nearly 136 million imaging examinations across seven US health systems and all of Ontario found that between 2000 and 2016, CT use in older adults rose from 204 to 428 exams per 1,000 person-years, and MRI from 62 to 139. Both roughly doubled. Growth slowed in later years but never reversed. Over a broadly overlapping period, the number of radiologists per capita in this country moved by less than half a radiologist per 100,000 people.

Growth in demand, growth in supply Each series indexed to 100 at its own starting year. Vertical scale starts at zero. 0 100 start end MRI, +124% older adults, 2000→2016 CT, +110% older adults, 2000→2016 Radiologists, +4% per 100,000, 2010→2022
Two curves that were never going to meet. The imaging figures come from Smith-Bindman's 2019 JAMA cohort (2000–2016); the workforce figure from a 2026 JACR analysis (2010–2022). The windows do not match, so read this as two separate trends placed side by side rather than a single like-for-like comparison. It is a picture of direction, not a calculation.

And the people absorbing that gap are not doing well. In a survey of a large coalition of physician-owned US radiology practices, 46% of radiologists met criteria for burnout and only 27% reported professional fulfillment. Taking call was the strongest associated factor. That survey had a 20.6% response rate, which means burned-out people may have been more motivated to answer, so treat the exact number loosely. The direction is not in dispute by anyone who works in a reading room.

In England, a 2026 review reported a 30% shortfall in clinical radiologists, projected to reach 40% by 2028.

This is the actual problem. Not "can a machine see the lesion." It is that scans are sitting unread, MRI slots are months out, and the people reading are running on empty. If you frame AI as a contest for who is the better reader, you have not even engaged with the thing that is hurting patients.

And it is not only radiologists

Step back from imaging and the picture gets genuinely strange, because the thing everyone is afraid of is not the thing that is happening.

The staffing crisis in American health care is not approaching. It is here, it has been here for a decade, and it keeps getting worse. In the ASRT's 2025 staffing survey, the CT technologist vacancy rate reached an all-time high of 19.4%. MRI was 17.4%. Cardiovascular interventional technology, 17.4%. Every imaging discipline surveyed sat above its 2020 level. Two years earlier the radiographer vacancy rate had hit 18.1%, up from 6.2% in 2021. The survey drew 475 department managers, so hold the decimal points loosely, but nobody who runs a department needs a survey to know this.

Unfilled positions being actively recruited, 2025 Share of posts vacant, by imaging discipline CT 19.4% MRI 17.4% Cardiovascular interventional 17.4% Nuclear medicine 12.6% Sonography 12.4% Mammography 11.4% 0% 20% CT is at an all-time high, up from 17.7% two years earlier, while CT volume has roughly doubled since 2000.
The jobs are open right now. Bars are zero-based and share one scale. Every discipline in this survey sits above its 2020 rate. From the ASRT 2025 Radiologic Sciences Staffing and Workplace Survey, which collected responses from 475 US radiology department managers. That is a small sample, so treat the decimal points as indicative rather than precise.

Now read that next to the volume figures from earlier in this post. CT use in older adults doubled. The CT technologist vacancy rate is nearly one in five. Those two facts together describe a queue, and the queue is made of people.

Then look at what the government actually projects, from the same 2025–2035 release I have been quoting on cashiers. Private health care and social assistance is projected to add more than 2.2 million jobs, the most of any sector, accounting for roughly 37% of all new jobs in the entire economy. Healthcare support and healthcare practitioners are the two fastest-growing of all 22 major occupational groups, and together they are expected to supply almost a third of every new job created through 2035. Nurse practitioner is the single fastest-growing detailed occupation in the country, at 41%.

In that same release, the group projected to shrink fastest is office and administrative support, shedding 752,100 jobs, the largest decline of any major occupational group. The BLS attributes it in plain language to automation, including AI.

Projected change in jobs, 2025 to 2035 Both figures from the same BLS release, published August 2026 Office and administrative support largest decline of any occupational group −752,100 Health care and social assistance about 37% of all new jobs in the economy +2,200,000 no change
Same document. Same decade. Opposite directions. Bars are drawn to a common scale from zero, so their lengths are directly comparable. The health care figure is reported by BLS as "more than 2.2 million," so the bar is a floor rather than an exact value. Source: BLS Employment Projections, 2025–2035, released 27 August 2026.

One government document, one projection round: office administration sheds three quarters of a million jobs while health care absorbs a third of all the new ones. That is not a story about work disappearing. It is a story about work moving.

I find the public conversation about this disorienting. I read that AI is about to leave people without jobs, and then I go to work in an industry that cannot fill the jobs it already has, where those unfilled positions are the direct reason somebody waits four months for a scan, and where the shortage has been deepening steadily since before large language models existed. The openings are posted. Where are the people?

So the question is not whether there is work. There is an enormous amount of work. The question is whether we are willing to move people toward it.

That requires being honest about which way things travel. Some roles genuinely are easier to automate, and cashiering is the clean case: what the customer needs is a fast, accurate, low-friction transaction, and a machine now delivers that well. Other roles are close to unautomatable on any near horizon, and they are disproportionately in health care. Turning a patient. Getting a difficult IV. Positioning someone who is in pain for a scan without hurting them more. Noticing that a person is frightened and doing something about it. Those are not knowledge tasks with a hands-on component. They are hands-on tasks with a knowledge component, and the order matters.

None of the moving happens by itself. Somebody has to fund the training pipelines, build bridges out of declining occupations into growing ones, and pay for the years in between. Enrollment in radiologic technology programs has been falling while the vacancies climb, which tells you the market signal is not reaching the people who could act on it. A labor market does not clear just because an economist can see that it ought to.

This is the part I would most like people to hear, because the fear is real and it is aimed at the wrong target. The jobs are not vanishing. They are relocating, into work that is harder to automate and, in most cases, more worth doing. What we owe people is the ladder to get there.

What this post does not tell you

Everything above is an argument about demand and labor. It is the argument I most want people to hear, because the fear is real and it is pointed at the wrong target. But it is also, deliberately, an argument that dodges the hardest question.

None of it tells you whether the technology actually works.

A staffing crisis is a reason to want a tool. It is not evidence that the tool is any good, and "we are desperate" is close to the worst frame available for a purchasing decision. There are real randomized trials now, some of them genuinely impressive. There is also a study in which very experienced radiologists went from scoring 82% of mammograms correctly to 46%. The difference was that the AI handed them the wrong answer.

That is the next post.


How this piece was built

I structured this on Randy Olson's And, But, Therefore framework, which keeps an argument from collapsing into a list of statistics. Every figure here was checked against the primary source, and where a survey is small or two datasets are not strictly comparable, I have said so in the text rather than hiding it in a footnote.

AI disclosure. The argument, the experience, and the point of view are mine. I worked with Claude (Anthropic) as a research and drafting partner: it verified every number against BLS tables, ASRT survey releases, and peer-reviewed sources, produced the four figures, and helped organize the draft. It corrected two things I had wrong going in, which I have left visible in the text: the 2016 prediction was Geoffrey Hinton's rather than Jensen Huang's, and my cashier figures were from a superseded BLS projection round. I reviewed every claim and citation before publishing.

References

  1. US Bureau of Labor Statistics. Employment projections: 2025–2035 summary. USDL-26-1422, August 27, 2026. bls.gov
  2. US Bureau of Labor Statistics. Occupations with the largest job declines, 2025 and projected 2035. bls.gov
  3. US Bureau of Labor Statistics. Occupational Employment and Wage Statistics, largest occupations, May 2019. bls.gov
  4. Bessen J. Toil and technology. Finance & Development (IMF), March 2015. imf.org
  5. American Society of Radiologic Technologists. 2025 Radiologic Sciences Staffing and Workplace Survey. asrt.org
  6. Smith-Bindman R, et al. Trends in use of medical imaging in US health care systems and in Ontario, Canada, 2000-2016. JAMA. 2019;322(9):843-856. doi:10.1001/jama.2019.11456
  7. Malhotra A, et al. The evolving US radiologist pipeline: trends in residency positions, resident workforce, and practicing radiologists per unit population. J Am Coll Radiol. 2026;23(8):1587-1592. doi:10.1016/j.jacr.2026.04.005
  8. Parikh JR, et al. Prevalence of burnout of radiologists in private practice. J Am Coll Radiol. 2023;20(7):712-718. doi:10.1016/j.jacr.2023.01.007
  9. Spalding A. A retrospective mixed-methods service evaluation of radiographer-led adult nephrostomy exchange service. Radiography. 2026;32(4S1):103316. doi:10.1016/j.radi.2025.103316

Peer-reviewed sources were located through PubMed and Consensus. Nothing in this post describes any individual patient or protected health information, and nothing here represents the position of my employer.

Tuesday, September 1, 2026

When the Hospital Came Home

My mother needed hospital care, and she got it. But the hospital also took her sleep, her movement, and every last bit of control over her own day. Then we qualified for a program that sent the hospital to her house instead, and almost everything except the medicine changed.

I am a radiologist. I have spent years inside hospitals, reading images for patients I rarely meet, trusting that the system on the other end of my report works the way it is supposed to. Then my mother was admitted, and I found out what the other side of that system feels like when you are the one sitting in the chair.

She does not speak English. That single fact reorganized our whole family. One of us had to be in the room with her at all times, so my siblings and I built a rotation and lived inside it. The room itself was lovely. Big windows, warm light, a recliner that folded back into something the brochure would probably call a bed.

It was not a bed. I know, because I spent nights in it.

The part nobody warns you about

Here is what I did not expect: the exhausting part was not the worry. It was the interruptions.

Vitals at midnight. A blood draw before dawn. An IV pump alarming at two in the morning. Someone coming in for weights, then someone else for the morning labs. My mother never got a full night. Neither did I. We were both awake at 4 a.m. in a room designed to make sure nothing was ever missed, which also meant nothing was ever quiet.

And then daylight brought the other problem: waiting. Several subspecialists were consulting on her case, and I never knew when any of them would appear. Rounds happened sometime. The nurse came sometime. The doctor will see you now, except no one could tell you when now was going to be.

So I did not leave. I skipped meals and held it and stayed put, because stepping out for ten minutes meant possibly missing the one conversation I had been waiting fourteen hours to have. I could not pick up my daughter. I could not be in two places. I sat in a beautiful room and felt completely trapped in it.

Meanwhile my mother sat too. Bed, chair, bed. Day after day, a woman who runs her own household barely moved twenty feet.

The care was excellent. The experience was not. Those turn out to be two different things, and the difference has a name in the literature.

The hospital itself is a stressor

In 2013, the Yale cardiologist Harlan Krumholz gave this a name in the New England Journal of Medicine: post-hospital syndrome. His argument is that the month after a hospital stay carries a broad, elevated risk of getting sick again, and much of that risk comes not from the original illness but from what the hospital did to the person while treating it. Sleep gets shredded. Nutrition suffers. People stop walking. Days lose their edges.

Once I read that, everything I had watched in that room stopped feeling like bad luck and started looking like a predictable pattern. The numbers back it up.

47 min
Less sleep per night in the hospital than the same patients got at home
683 older medical inpatients, four hospitals (Smichenko 2025)
57%
Of observed daytime hours, inpatients of all ages spent lying in bed. Nine percent standing or walking.
132 inpatients, behavioral mapping (Mudge 2016)
30%
Of hospitalized older adults go home less able to do a basic daily task than when they arrived
Meta-analysis, 7,375 patients (Loyd 2019)

That last one has a clinical name too: hospital-associated disability. Someone walks in able to bathe or dress themselves and walks out unable to, and the thing that took it away was the stay, not the illness. A large prospective study found that in-hospital mobility, continence care, and length of stay together explained 64% of the variation in who declined by discharge. Those are all things a system chooses.

Bed rest is often not even a medical decision. In a study of 498 hospitalized adults over 70, a third had bed rest ordered at some point, and among the least mobile patients, nearly 60% of those bed rest episodes had no documented medical reason at all. We immobilize people out of habit.

The sleep piece is just as fixable and just as stuck. When researchers asked patients, physicians, and nurses what wrecks sleep in the hospital, all three groups named the same top three: pain, vital signs, and tests. Everyone knows. It happens anyway.

And the language problem sitting underneath all of it

My family's rotation existed because my mother could not advocate for herself in English. I used to think of that as our private logistics problem. It is actually a documented safety issue.

In a study of 1,666 families across seven North American hospitals, children whose parents were not comfortable speaking English in medical settings had roughly twice the odds of experiencing a harm caused by their medical care (17.7% versus 9.6%). A companion study across 21 hospitals found families with limited English proficiency were dramatically less likely to speak up when something looked wrong, or to question a clinician's decision. Both studies looked at hospitalized children rather than adults, so I hold the specific numbers loosely. The direction is not in doubt, and it matched our experience exactly. Being physically present was our workaround for a system that could not hear her.

Then we qualified for hospital at home

Hospital at home is not new. Versions of it have run for decades in Australia and the United Kingdom. What is new in the United States is scale, and the reason is technology plus a Medicare waiver. The model works like this: a physician-led team runs your care from a command center, in-person visits come to your house, and everything that can be done remotely is. You are formally an inpatient. You are just an inpatient in your kitchen.

A program sets you up at home with a tablet, a blood pressure cuff, a scale, an oxygen monitor, an emergency alert you wear, a direct-dial phone to the command center, plus a wifi extender and a backup power supply because the whole thing depends on staying connected. Physicians, nurse practitioners, pharmacists, nurses, social workers, and paramedics all work off the same plan.

My mother qualified. Here is the day that followed.

In the hospital At home, still an inpatient midnight 3 am 6 am 9 am noon 3 pm 6 pm 9 pm Vitals IV pump alarm Vitals Labs drawn Weights, shift change Rounds. Sometime. Do not leave the room. Do not shower. Do not go get lunch. Consultant, unannounced Consultant, unannounced Vitals Vitals Asleep. Nobody comes in. Button within reach if anything changes. I slept in my own bed. Nurse video visit, 8:00 Labs, scheduled window Doctor visit, time we picked Lunch at the table. Laundry. Dishes. Walking. Nurse video visit Paramedic, in person Nurse video visit Lights out
Same illness, same medicine, two different days. A composite of our experience, not a chart of measured data. The hatched blocks on the left are the part that wore me down: care that was definitely coming, at a time nobody could tell me.

What actually changed

My mother started moving. Not because anyone prescribed it, but because she was in her own house and there was laundry to fold and dishes in the sink. She got up. She walked around. She did her own things. Within a day she was doing more than she had done in a week of lying in a beautiful room.

I slept. Fully, in my own bed, and nobody came in at 2 a.m. I picked my daughter up. I was in the same building as both the person I care for and the person I am raising, which had felt impossible for weeks.

And the schedule became ours. The nurse came at a time we knew. The blood draw had a window. I could actually schedule my mother's visit with the primary team, which meant I could plan a day around it instead of surrendering the day to it. Consultants still appeared without warning, but they appeared on a screen for a few minutes rather than being an eight-hour vigil.

There was one more thing I did not anticipate. In the hospital, I felt guilty calling the nurse. Towels, another gown, a cup of coffee: I knew how busy she was and I could see her running, so I sat on small needs and let them stack up. At home, a nurse was one button away, twenty-four hours a day, and I used it without hesitation, because now every call I made was actually about my mother's care. The small stuff was just ours to handle. That reallocation felt better for everyone.

It was not only that the care moved. It was that we got the environment back. The medicine stayed the same and the power over the day came home with her.

I thought this was my private observation until I found a study that had written it down. Researchers interviewed patients from a randomized home hospital trial and found that home patients described "a locus of control surrounding their sleep, activity, and environmental comfort" that hospitalized patients simply did not have. That is the whole thing, in the dry language of qualitative research. Not comfort. Control.

Does it actually work, or does it just feel better?

This is where I put my radiologist hat back on, because a good feeling is not an outcome. The honest answer is that hospital at home holds up on safety and wins clearly on experience and activity, while the cost and readmission findings are real but less consistent than the enthusiasm suggests.

Traditional hospital Hospital care at home
Readmitted within 30 days Boston randomized trial, 91 patients 23% 7% Share of the day spent lying down Same trial, measured by accelerometer 55% 18% Felt “extremely” or “very” comfortable Randomized trial, 1,150 patients 60.9% 84.4%
Two separate randomized trials. Compare the two bars within a panel, never across panels. The Boston trial (Levine 2020) randomized only 91 highly selected patients at two sites, with 63% of eligible patients declining to participate, so treat those two panels as promising rather than settled. The comfort figure comes from a larger 2025 trial of 1,150 patients (Maniaci 2025).

The strongest single piece of evidence is a randomized trial published in 2025. It randomized 1,150 acutely ill patients across three hospitals to hospital-at-home care or a traditional bed. The combined rate of death or unplanned readmission within 30 days was 17.3% at home and 19.8% in the hospital, which met the trial's bar for showing home care is not worse. No patient died while receiving their hospital care at home. And on comfort, the gap was wide: 84.4% versus 60.9%.

The broader picture, from a Cochrane review of 20 randomized trials covering 3,100 people, is consistent. Hospital at home probably makes little or no difference to death rates or readmissions, probably lowers costs, and probably makes people meaningfully less likely to end up living in a nursing home six months later. That last finding deserves more attention than it gets.

Two honest caveats. First, cost savings are not automatic: when Levine's group ran the same model in rural communities in 2025, the episode cost came out no different from a regular hospital stay, even though patients took roughly seven times more steps per day and rated the experience far higher. Second, almost every one of these trials enrolled carefully selected, relatively stable patients. That is exactly who the program is for, and it is exactly why you cannot generalize the results to everyone in a hospital bed.

Where it is imperfect, including for us

I do not want to write a brochure. We had real friction.

The tablet needed rebooting. We had connection problems. When technology is the spine of your care, the spine occasionally goes out. What made it workable was that the program planned for exactly this: there were two separate backup ways to reach the team while the tablet was being sorted out. Redundancy is not a nice-to-have in this model, it is the safety system. If you are evaluating a program, ask what happens when the internet drops, and do not accept a vague answer.

The bigger caveat is the one the research keeps flagging and the marketing keeps skipping: this model leans on the family. In a study of 125 caregivers assessed in the first 48 hours of a hospital at home admission, 61.6% already met the threshold for high caregiver strain. Interviews with caregivers in the United States and Denmark found the same pattern: they overwhelmingly preferred it to a hospital stay, and they also felt underprepared, unclear about what was their job versus the team's job, and sometimes overwhelmed.

I had three siblings, a flexible enough job, and clinical training. That is not most families. A model that quietly assumes a capable, available caregiver will work beautifully for people who have one and will not be offered to people who do not. That is an equity problem sitting right in the middle of a very good idea, and it should be designed for rather than discovered later.

If a program is offered to your family, ask these

  • Exactly what am I responsible for, and what is the team responsible for? Get it in writing.
  • What are the backup ways to reach you if the tablet or the internet fails?
  • How fast can someone physically get to the house, and who is that person?
  • Can we schedule the daily physician visit, or does it just happen?
  • What triggers a transfer back to the hospital, and how does that work at 3 a.m.?
  • Is interpretation built into every visit, or does the family have to arrange it?

Why I am excited about this beyond my own family

For twenty years, nearly every effort to improve the patient experience has aimed at making the in-person visit better. Nicer rooms. Better food. Softer lighting. My mother's room proved the ceiling on that strategy: it was a genuinely beautiful room, and it still took her sleep, her movement, and her control, because those losses are structural rather than decorative.

Remote monitoring plus a command center does something a renovation cannot. It keeps the medicine and drops the institution.

And it frees a bed. Australia's Victorian hospital-in-the-home program was described in one paper as "the 500-bed hospital that isn't there." Every stable patient treated at home is a bed available to someone who is critically ill and genuinely needs hands on them, in a country where capacity is the binding constraint on almost everything. My mother going home was not just better for my mother. It was better for whoever got that room.

The policy question is settled for now. Congress extended the Medicare waiver through 2030 as part of the Consolidated Appropriations Act, 2026. As of that extension, 366 programs across 139 health systems in 37 states were approved to deliver acute hospital care at home. Five years of stability is enough runway for health systems to actually build rather than pilot.

Framed against the Quintuple Aim, the goals most of us in health care now organize around, this model plausibly moves four of the five at once: outcomes hold, experience improves substantially, costs trend down in most settings, and freed capacity helps the sickest patients. The fifth, equity, is the one that will not take care of itself. It depends entirely on whether programs get built for families who do not already have a spare adult and a strong wifi signal.

What I keep coming back to

My mother received the same medicine either way. Same labs, same monitoring, same physicians. What changed was that she got to fold her own laundry, sleep through the night, and eat lunch at her own table, and I got to be a daughter and a mother on the same day instead of choosing.

For years I assumed the goal was a better hospital. I think the actual goal is needing the hospital for less.


How this piece was built

I structured the story on Randy Olson's And, But, Therefore framework, which is a simple way to keep a narrative from collapsing into a list of facts. Three published frameworks shaped how I read my own experience: Krumholz's post-hospital syndrome for why the hospital itself is a stressor; the Age-Friendly Health Systems 4Ms (What Matters, Medication, Mentation, Mobility) for naming what changed at home; and the four core concepts of patient- and family-centered care (respect and dignity, information sharing, participation, collaboration), which is where the agency argument actually lives.

AI disclosure. I wrote this from my own experience and my own point of view. I worked with Claude (Anthropic) as a research and drafting partner: it searched PubMed and Consensus for the peer-reviewed evidence cited here, verified the trial numbers against the source abstracts, identified the narrative and conceptual frameworks above, produced the two figures, and helped me organize and tighten the draft. I directed the argument, supplied the experience, and reviewed every claim and citation before publishing.

References

  1. Maniaci MJ, et al. Safety in a hybrid hospital-at-home program versus traditional inpatient care: a pragmatic randomized controlled trial. J Hosp Med. 2025;20(11):1174-1184. doi:10.1002/jhm.70076
  2. Levine DM, et al. Hospital-level care at home for acutely ill adults: a randomized controlled trial. Ann Intern Med. 2020;172(2):77-85. doi:10.7326/M19-0600
  3. Levine DM, et al. Hospital-level care at home for adults living in rural settings. JAMA Netw Open. 2025;8(12):e2545712. doi:10.1001/jamanetworkopen.2025.45712
  4. Levine DM, et al. Hospital-level care at home for acutely ill adults: a qualitative evaluation of a randomized controlled trial. J Gen Intern Med. 2021;36(7):1965-1973. doi:10.1007/s11606-020-06416-7
  5. Edgar K, et al. Admission avoidance hospital at home. Cochrane Database Syst Rev. 2024;3(3):CD007491. doi:10.1002/14651858.CD007491.pub3
  6. Krumholz HM. Post-hospital syndrome: an acquired, transient condition of generalized risk. N Engl J Med. 2013;368(2):100-102. doi:10.1056/NEJMp1212324
  7. Loyd C, et al. Prevalence of hospital-associated disability in older adults: a meta-analysis. J Am Med Dir Assoc. 2020;21(4):455-461.e5. doi:10.1016/j.jamda.2019.09.015
  8. Brown CJ, et al. Prevalence and outcomes of low mobility in hospitalized older patients. J Am Geriatr Soc. 2004;52(8):1263-1270. doi:10.1111/j.1532-5415.2004.52354.x
  9. Zisberg A, et al. Hospital-associated functional decline: the role of hospitalization processes beyond individual risk factors. J Am Geriatr Soc. 2015;63(1):55-62. doi:10.1111/jgs.13193
  10. Mudge AM, et al. Poor mobility in hospitalized adults of all ages. J Hosp Med. 2016;11(4):289-291. doi:10.1002/jhm.2536
  11. Grossman MN, et al. Awakenings? Patient and hospital staff perceptions of nighttime disruptions and their effect on patient sleep. J Clin Sleep Med. 2017;13(2):301-306. doi:10.5664/jcsm.6468
  12. Smichenko J, et al. Sleep trajectory of hospitalized medically ill older adults. Sleep. 2025;48(5):zsaf013. doi:10.1093/sleep/zsaf013
  13. Khan A, et al. Association between parent comfort with English and adverse events among hospitalized children. JAMA Pediatr. 2020;174(12):e203215. doi:10.1001/jamapediatrics.2020.3215
  14. Khan A, et al. Association of patient and family reports of hospital safety climate with language proficiency in the US. JAMA Pediatr. 2022;176(8):776-786. doi:10.1001/jamapediatrics.2022.1831
  15. Duhamel S, et al. Caregiver burden at the onset of acute hospital-at-home. J Am Geriatr Soc. 2026;74(8):2338-2348. doi:10.1111/jgs.70573
  16. Bertelsen KB, et al. When the home becomes the setting for hospital treatment: a qualitative study of relatives' experiences. J Adv Nurs. 2025;82(1):567-579. doi:10.1111/jan.16955
  17. Montalto M. The 500-bed hospital that isn't there: the Victorian Department of Health review of the Hospital in the Home program. Med J Aust. 2010;193(10):598-601. PMID 21077817
  18. Nundy S, Cooper LA, Mate KS. The Quintuple Aim for health care improvement. JAMA. 2022;327(6):521-522. doi:10.1001/jama.2021.25181

Peer-reviewed sources were located through PubMed and Consensus. Nothing in this post describes anyone's diagnosis or protected health information.