Most conversations about AI in radiology ask whether it will read the images. That has never been the part of my job that consumes me. What consumes me is everything stacked around the images: the report I have to reconcile against six prior studies, the lecture that needs twelve good cases pulled from ten years of archives, the call schedule that has to satisfy a dozen rules, the license I have to renew on a website designed by someone who has never renewed one.
That is where AI has changed my work. It does not read for me. It clears the ground so I can read.
I want to describe what that actually looks like, because most of what I read about AI in medicine is either breathless or dismissive, and neither matches my experience. Mine has been slower and more mundane than the hype, and more useful than the skepticism.
One note before I start. Everything below runs inside enterprise, HIPAA-covered instances of these tools that my institution has approved, and the research is IRB-approved. None of this involves pasting patient information into a consumer chatbot, and I would not advise anyone to do that.
Reporting: the workflow I measured
Every one of my reports now goes through a custom GPT.
I built a set of them in ChatGPT Enterprise, one per modality, each with disease-specific templates. I dictate into the ChatGPT dictation button rather than into PowerScribe. The dictation is simply better, and that is the whole reason I switched. It is the one thing I use ChatGPT for exclusively, and everything else sits on top of it.
Three of these have become indispensable.
The "MISS" template. I paste in all the prior radiology reports along with my current draft, and it checks my current report against the priors to confirm I have addressed every finding someone described before. This one earns its keep on oncology studies, where a patient may have eight prior scans and a dozen tracked lesions. What goes wrong there is not that I misread something. It is that a finding described eighteen months ago quietly drops out.
The revised-report template. I review the prior findings, tell it what to change, and it writes a new draft.
Fat fraction. I paste in the in- and opposed-phase images with ROIs over liver and spleen, and it hands back a calculation formatted and ready to drop into the report. No calculator, no retyping.
I studied this rather than just claiming it worked. I published the results in Abdominal Radiology, framed through the Unified Theory of Acceptance and Use of Technology, comparing my own baseline period against my post-implementation period across 609 studies.
The findings split cleanly. For outpatient CT, the gains were large. With contrast, my median inter-study interval dropped from 23 minutes to 13. Without contrast, it dropped from 18.5 minutes to 7. Both were significant. For MRI, nothing clear. With contrast the numbers drifted slightly the wrong way, and without contrast the improvement looked big but did not survive correction for multiple comparisons.
I think that null result is the most honest and most useful thing in the paper. Standardized, high-volume CT is exactly the task a templated LLM workflow fits. Complex MRI is not. The cases vary more, the templates are harder to build, and you cannot reduce the reasoning to a form. How well the tool fits the task is doing the real work here, not the model.
Training took ten hours across five days. That is not nothing, and it is worth saying plainly, because people underestimate what you have to put in before you get anything back.
I have also looked at what LLMs do to report accuracy, in a separate study where six radiologists reviewed GPT-4's suggested revisions to 600 of their own finalized abdominopelvic CT reports. GPT-4 flagged something in 91% of reports, but the radiologists accepted only 23% of what it suggested, and most of what it caught was grammar. The clinically meaningful catches were real but uncommon.
I raise that because it tempers my own enthusiasm. The reporting workflow makes me faster and gives me a systematic second pass against priors. It does not make me infallible, and the data do not support anyone who says otherwise, including me.
Teaching: an assistant that does not go home
I am putting together a prostate case review talk for SABI 2026 in Savannah. Historically this takes weeks. Think of the teaching points, hunt through the archive for cases that actually demonstrate them, then chase down the pathology and the clinical follow-up.
Now I brainstorm the case list with Codex, then point it at our enterprise Illuminate instance, which holds the radiology, pathology, and clinical notes together. It pulls the cases while I look them up on PACS and grab images for the slides.
The moment it clicked for me was small. Codex found me a Müllerian duct cyst, but it measured 1.6 cm, too subtle to teach from. I asked for something bigger. It came back with a 15 cm cyst, now too dramatic to represent anything. I asked for something in between and got it.
That exchange took under a minute, and I would never have asked a human assistant to redo the same task three times in a row. It feels like working alongside a superintelligent assistant who does not get tired of my revisions.
Scheduling: the work nobody wants
I run the schedule for the Abdominal Radiology Case Conference, and scheduling has always been the most thankless part of it. Email everyone. Remind them to submit their availability. Reconcile what everyone gives you against the rules. Send calendar invites. Chase the people who never replied.
A student intern and an administrative assistant used to handle it, which worked until someone traveled or a new person came on, and then it broke and left gaps.
Claude now runs the whole cycle: the outreach, the reminders, the rule-based matching, the invites, the follow-ups. So far it is working well.
Our clinical schedule in QGenda got the same treatment. The QGenda rules are not hard, they are just tedious, which is exactly the kind of task worth handing off. I built a Codex project, gave it all the definitions and rules, told it what I wanted, and it took care of it. I have since found I prefer Claude for this particular job, because it gives me a visual dashboard and tracks my schedule as it changes rather than just answering once.
A smaller one in the same vein. I attend a weekly educational meeting where sessions sometimes get cancelled, the instructor changes, and each instructor sends a different Zoom link. All of it arrives by email. Now I hand Claude the email and tell it to update my calendar with the right link and the right instructions, so what is on my calendar always matches what is actually happening that week.
Research: the long tail
I have a project running now on discordant prostate MRI, meaning patients whose MRI was positive but whose biopsy came back negative, across roughly 500 patients. Codex is working through the follow-up MRIs, repeat biopsies, notes, and labs.
I layer the tools, and the layering is the point. I use enterprise Claude to write the instructions for Codex, then to troubleshoot when Codex gets stuck and to double-check what it produces. Having one model brief and audit another has caught things neither would have caught alone.
The small things, which turn out not to be small
I renewed my medical license recently. Anyone who has done it knows the actual work is trivial and navigating the site is miserable. You hunt for the right link, the right page, the right form, before you can start. Codex walked me through it and made it painless.
I mention this because it is typical. So much of my week goes to digging for the right place to begin rather than doing the work itself. Clearing that friction has changed how my days feel more than any single clinical application has.
The same holds for my student interns. I get a new one every year, and retraining them eats a lot of time. The workflows now absorb most of the repetitive tasks, so the time I spend with an intern goes toward actually teaching them something.
And announcing our division's recent publications on social media now runs automatically. Small thing. Off my plate.
Caregiving: the one I would least want to give back
I also care for an aging parent who is ill, which means a steady stream of appointments and a family calendar that has to stay in sync around them. I log into the patient portal, and Claude takes every appointment and puts it on our family Google Calendar, then keeps it updated as things shift.
None of that is hard. It is tedious, it arrives in fragments, and it never really stops. Handing it off lifted a specific kind of mental weight I had stopped noticing I was carrying. Of everything on this list, this is the one I would least want to give back.
What it gave back
Two effects I did not expect.
The reports got better, not just faster. Speed is what I set out to measure. Quality is what I noticed afterward. A systematic second pass against every prior raises the floor on a heavy day, which is exactly the day a finding from eighteen months ago slips through.
I stopped falling behind. The tasks I used to carry around as low-grade dread, the ones that were never hard but were always waiting, are now automated or scheduled. That dread took up more room in my head than the tasks ever took in my week.
What surprised me is what filled the space. I have wanted to write here for years and never had the bandwidth. This post exists because the scheduling, the calendar updates, and the reminders stopped eating the hours I would have spent on it.
It spilled into my personal life too. I have friends who fly first class and take their families on essentially free vacations using credit card points, and they have been telling me about it for years. I always found it interesting in theory, and I was never going to sit down and learn it.
So I built a Claude project instead. I gave it the collective knowledge from the physician points community I follow, the transfer rules between programs and their partners, my credit cards, all of my points accounts, and my travel goals for 2027, and I let it work.
It went through my cards and told me I could drop my Sapphire Reserve and save roughly $800 a year, but that I should downgrade it to a Freedom card rather than cancel it. Downgrading keeps my Ultimate Rewards points alive and avoids putting a closed account on my credit history. I had no idea downgrading was even an option. It also showed me that my other premium cards already carry most of the benefits I was paying the Reserve for, so I had been paying twice for the same thing.
That is one example out of many. The pattern is the same as it is at work. The barrier was never that any of this was difficult. The barrier was that I was never going to make the time.
What I have actually learned
Each tool is good at something different, and it is worth finding out what. ChatGPT dictates best, so all my reports go through it. Codex is where I work across systems. Claude gives me better visual, structured output, which is why my schedule lives there. I did not decide this in advance. I found it by using all three badly for a while.
You have to invest a lot before you get anything back. Building these workflows took hours I did not obviously have, and I still spend many hours a day working with AI. But once a workflow is built, it runs close to seamlessly, and it pays off a little more every week.
Fit matters more than raw capability. My own data showed a large benefit for CT and none for MRI, with the same model and the same radiologist. What differed was the task, not the technology. I would rather build three workflows that fit than ten that impress.
Measure it. I have written before that you cannot improve what you cannot measure, and I meant it about teaching. It applies here too. It would have been easy to feel faster and never check. Some of what I believed held up. Some of it did not.
None of this replaced my judgment. It cleared away the things standing between me and the point where judgment is required. I still read every image and I still own every report. I just spend a much larger share of my day on the part that actually needs me.
If you want to start
A few things I would tell someone at the beginning.
Tell it to double-check its work. Every time. This is the highest-yield instruction I give, and it costs one sentence.
Use two. I keep both Claude and ChatGPT, personal and enterprise. Ask them the same question and you get different answers, and the difference is the useful part. It is complementary rather than duplicative, like having two very smart consultants who think differently. I often use one to audit the other.
Watch a short video. I got started with Codex from one 28-minute YouTube video: Learn 95% of Codex in 30 minutes by Riley Brown. That was the whole onboarding.
You do not need to code. I am not a coder. Nothing I described in this post required me to be one.
Start small. You are not going to build Rome overnight. Every workflow here began as one annoying task I decided to hand off, and they accumulated from there.
One thought I keep returning to. I am starting to want a different kind of student intern: someone who supervises and runs AI workflows rather than doing every task by hand. But that only works if the person still understands the workflow deeply. You cannot supervise a process you do not understand, and you cannot tell when the output is wrong if you have never done the work yourself. That is the part that does not get automated.
How I used AI to make this post. Claude and I co-wrote this post. I talked through my workflows in one long unedited brain dump, and Claude turned that into the structure, the section order, and the prose you just read. Claude searched PubMed for my own papers, pulled the exact figures out of them, and added a fourth reference I had forgotten I was an author on. It built all three data figures here from the published numbers. It also fact-checked me: it cut a claim I made that my reports contain no mistakes, because my own GPT-4 paper does not support that, and it caught a broken link to our case conference channel before this went live. I gave it my writing preferences and it applied them. I read every line, changed what I wanted changed, and approved the final version. The experience and the opinions are mine.
Opinions are my own.
References
- Tan N. Large language model-assisted radiology reporting in a single-radiologist implementation: a retrospective cohort study interpreted through a UTAUT lens. Abdominal Radiology (NY). 2026. doi:10.1007/s00261-026-05524-y
- Mayes CJ, Reyes C, Truman ME, et al. Improving radiology reporting accuracy: use of GPT-4 to reduce errors in reports. Abdominal Radiology (NY). 2025;51(1):513-520. doi:10.1007/s00261-025-05079-4
- Reyes C, Nguyen E, Alexander LF, et al. Beyond Human Limits: The Promise and Pitfalls of Large Language Models in Radiology Research. Journal of Computer Assisted Tomography. 2025;49(4):545-553. doi:10.1097/RCT.0000000000001709
- Gerbaud AF, Berguido de la Guardia M, Booth SC, et al. Process Improvement Before Artificial Intelligence and Automation: Building Trust With the Understand-Transform-Sustain Framework. Mayo Clinic Proceedings: Innovations, Quality & Outcomes. 2026;10(4):100725. doi:10.1016/j.mayocpiqo.2026.100725
No comments:
Post a Comment