AI in Medicine

AI in Medicine Review: 6 Stories From August-September 2026

A physician's review of six AI in medicine stories: ChatGPT in Epic, FDA outcomes evidence, SkinVision, AI ultrasound, autonomous mammography triage, and oncology RAG summaries.

By Mendel Jacobs, MD·Sep 9, 2026·12 min read
Physician using ChatGPT for healthcare to summarize clinical notes, labs, medications, and medical evidence

August and the first week of September gave us a useful snapshot of where AI in medicine actually is.

Not where the demo videos say it is. Not where the LinkedIn posts say it is. Where it is.

The optimistic version is real: AI-guided ultrasound may help non-experts screen for aortic stenosis, retrieval-augmented generation can pull key oncology facts out of messy records, and autonomous imaging workflows are starting to move from theory into regulation.

The uncomfortable version is also real: a skin-cancer app can look cheaper partly because it detects fewer cancers, and FDA-authorized medical AI still has very little prospective patient-outcomes evidence behind it.

That is the theme this month: AI is getting closer to the clinical workflow faster than the evidence base is getting closer to patients.

The short version

Here is how I would file the six stories.

StoryWhat happenedMy read
ChatGPT and the EHROpenAI announced healthcare source integrations, including a read-only Epic connection being piloted at UCSF Health.Integration is the product. The open question is whether it saves time without missing the thing you needed.
FDA AI evidence gapA PLOS Digital Health census found 1,357 FDA-authorized AI/ML medical devices but only 3 with prospective patient-centered outcomes evidence.FDA-authorized is not the same thing as outcome-proven.
SkinVisionA Belgian cost-effectiveness model found lower costs but substantially lower cancer detection compared with usual care.Cheaper is not better if the pathway is cheaper because it finds less disease.
AI-guided ultrasoundNovices with four hours of training used AI-guided focused cardiac ultrasound to screen for aortic stenosis.This is the kind of access-expanding AI that feels clinically plausible.
Autonomous mammography triageVara received European certification for autonomous triage of clearly normal screening mammograms.Autonomy makes monitoring, fallback, and drift detection part of the medical device.
Oncology RAG summariesAn MSK study found AI-generated breast-oncology summaries were better on several structured facts but had more clinically significant fabrications.RAG can reduce chart review pain. It can also create very polished chart lore.

1. ChatGPT is getting closer to the EHR

OpenAI announced a healthcare version of ChatGPT that can connect to medical sources, including a read-only Epic integration that UCSF Health is piloting.[1]

The pitch is straightforward: authorized clinicians can pull in notes, labs, medications, and other chart information, then ask ChatGPT to summarize, organize, or surface relevant context. OpenAI also described a healthcare public-data connector that links ChatGPT to sources like PubMed, ClinicalTrials.gov, DailyMed, and CMS coverage information.[1]

This is not ChatGPT practicing medicine. It is ChatGPT trying to sit where clinicians already work.

That distinction matters. A general chatbot in a separate tab is interesting. A model with access to the EHR, research databases, and the workflow where clinicians already spend their day is different. The distribution problem starts to shrink.

OpenAI also reported physician safety evaluations across 27 clinical use cases, with 99.1% of responses rated safe.[1] That is worth noting, but I would keep the champagne in the fridge. Company-run safety evaluation is a start, not a patient-outcomes trial.

The question I care about is more boring and more useful: does this make the clinician faster, more accurate, or less cognitively overloaded without quietly omitting the one fact that changes the plan?

That is the study I want to read.

2. FDA-authorized medical AI has an outcomes problem

The most important macro paper of the month came from Abulibdeh and colleagues in PLOS Digital Health.[2]

The investigators looked across the FDA's AI/ML-enabled medical-device database and asked a simple question: how many authorized products have actually been prospectively studied, published, and tested on outcomes patients care about?[2][3]

The numbers are not subtle.

  • 1,357 FDA-authorized AI/ML-enabled devices.
  • 34 devices, or 2.5%, had a linked registered prospective trial.
  • 12 devices, or 0.9%, had posted trial results.
  • 12 devices, or 0.9%, had peer-reviewed trial publications.
  • 3 devices, or 0.2%, evaluated patient-centered outcomes.

The authors used a strict and clinically meaningful definition of patient-centered outcomes: mortality, major morbidity, hospitalization or readmission, validated quality of life, function, or symptom burden. AUC did not count. Workflow speed did not count. A technically cleaner image did not count.

That is exactly the right provocation.

To be fair, this does not mean 99.8% of FDA-authorized AI products do not work. FDA clearance is not designed to require a mortality trial for every algorithm. A tool that measures anatomy on a CT scan should not have the same evidence burden as a tool recommending admission, treatment, or discharge.

But the paper exposes the gap between the language around medical AI and the evidence actually collected. If a company says a product is transforming care, improving outcomes, or saving lives, it is reasonable to ask whether any patient-centered outcome has been measured.

My rule: the evidence standard should scale with clinical authority. The closer the AI gets to diagnosis, triage, treatment, referral, or screening closure, the less acceptable it becomes to stop at retrospective accuracy.

3. SkinVision shows the danger of cheap-but-less-diagnostic AI

The SkinVision paper is a good reminder that cost-effectiveness can bite if you do not look carefully at what is being optimized.[4]

Meertens and colleagues modeled an AI-based smartphone app for early skin-cancer detection in Belgium, comparing the app pathway with usual care. The model looked at melanoma, basal-cell carcinoma, squamous-cell carcinoma, and benign lesions separately. The time horizon stopped at confirmed diagnosis, which matters.

The headline could sound encouraging: the AI pathway was cheaper.

But look at why.

ConditionAI app detectionUsual care detectionAI pathway costUsual care cost
Melanoma60.1%87.4%EUR 419EUR 536
Basal-cell carcinoma58.5%91.2%EUR 143EUR 145
Squamous-cell carcinoma70.6%92.2%EUR 228EUR 233

So yes, the app pathway cost less. But it also detected substantially fewer cancers. The authors concluded that, at current performance, the app did not support implementation or reimbursement in Belgium.[4]

This is not a trial showing the app harmed patients. It is a model, and because the time horizon ended at diagnosis, it did not fully capture the downstream cost of delayed cancer detection.

Still, the conceptual point is excellent. You can make a pathway cheaper by doing less of the thing the pathway was supposed to do.

That is not medical progress. That is accounting with a stethoscope.

4. AI-guided ultrasound for aortic stenosis is the optimistic version

The JAMA Cardiology aortic-stenosis paper is one of the strongest positive studies in this batch.[5]

The investigators did not just ask whether AI could interpret echocardiograms acquired by expert sonographers. They tried to close more of the loop:

  1. AI guides a novice operator to acquire the focused cardiac ultrasound images.
  2. AI interprets those images for aortic stenosis.
  3. Experts review only the subset most likely to need adjudication.

The model was developed using 6,753 patients. In the prospective component, novice operators performed 1,302 focused cardiac ultrasound exams after only four hours of training. Of those, 1,258 exams, or 96.6%, were suitable for automated analysis.[5]

For moderate-or-greater aortic stenosis, AI showed 93% sensitivity and 96% specificity. The raw positive predictive value was only 49.4%, which would be rough as a standalone diagnostic workflow.

But the second step made the design more interesting. When an expert reviewed only AI-positive and uninterpretable studies, roughly 10% of all exams, the positive predictive value rose to 91.1%. Sensitivity fell to 85.4%.[5]

That is a real tradeoff, but it is also a realistic workflow.

The value proposition is not "replace cardiology." It is "screen more people, in more places, and concentrate expert labor where it matters." That could be meaningful for rural clinics, primary-care offices, community screening, or health systems with limited echo capacity.

The remaining question is outcomes. Does this detect more previously undiagnosed aortic stenosis? Does it shorten time to valve intervention? Does it improve symptoms, hospitalization, or survival? We do not know yet.

But as a model of AI deployment, this is much more convincing than another leaderboard.

5. Europe authorized AI to clear normal mammograms without a radiologist

On September 2, 2026, Vara announced Class IIb CE certification under the European Medical Device Regulation for autonomous triage in organized breast-cancer screening.[6]

The workflow is simple to describe and very important: if the AI classifies a screening mammogram as clearly normal, no radiologist has to read that exam. Everything else still goes to a human reader.

That is materially different from most imaging AI. A lot of AI in radiology flags, prioritizes, or assists. This closes some cases.

The most interesting part is the monitoring architecture. Vara describes a continuous monitoring layer called ATMON that watches site-specific performance, hardware changes, system health, cancer-detection signals, and recall signals, with the ability to revert a site to full radiologist review if performance moves outside predefined limits.[6]

That is the right direction, because autonomy changes the validation question. For decision-support AI, the clinician is usually the explicit safety net. Once the AI can clear a case, monitoring and fallback are no longer optional software features. They are part of the medical device.

Vara's prior PRAIM study prospectively evaluated AI-supported mammography screening in more than 461,000 women.[7] But the newly certified autonomous workflow goes beyond the exact workflow directly tested in that trial.

So the next evidence question is not only whether the classifier performs well. It is whether the whole autonomous system performs well after deployment, site by site, scanner by scanner, month by month.

Autonomy is not a switch. It is a surveillance program.

6. Oncology RAG summaries were better and more dangerous

The Memorial Sloan Kettering breast-oncology paper may be my favorite generative-AI workflow study of the month.[8]

Moo and colleagues randomly selected 200 new breast surgical-oncology consultations. A retrieval-augmented large language model generated structured summaries from the relevant pathology and radiology reports. The investigators compared those AI summaries with the actual staff-authored preconsult documentation.[8]

The AI did very well on several discrete facts:

  • Correct nodal status: 95% with AI versus 72% with staff-authored documentation.
  • Presence of invasive tumor on biopsy: 94% with AI versus 30%.
  • Receptor status: 86% with AI versus 96% with humans.

Then comes the part that matters.

At least one fabricated item appeared in 8% of AI summaries versus 3% of staff summaries. Clinically significant fabrications occurred in 3% of AI summaries versus 0.5% of staff summaries.[8]

That is the generative-AI problem in miniature. The same tool can be more complete, more accurate on several structured variables, and more likely to invent a clinically meaningful detail.

The likely future workflow is not "AI writes the chart and the clinician goes home early." It is more like: AI extracts the tedious facts, then the clinician verifies the high-risk variables before the visit.

That could still be valuable. Chart review is real work. Previsit synthesis is real work. But the next trial should measure total clinician time, verification burden, error rates, cognitive load, downstream decisions, and patient outcomes.

Until then, I would call this promising and absolutely not self-driving.

What I am taking from this month

I keep coming back to four practical points.

First, integration is beating cleverness. The tools that matter are moving into the EHR, imaging workflow, previsit workflow, and screening program. A better model in the wrong place is still a tab nobody opens.

Second, the evidence bar has to match the clinical risk. A documentation draft, a measurement tool, a screening triage system, and a treatment recommendation tool should not face identical evidence standards.

Third, "human in the loop" is not magic. It only works if the human knows what to verify, has time to verify it, and is shown the uncertainty clearly. Otherwise it becomes a liability phrase wearing a badge.

Fourth, citations and RAG are not a force field. Retrieval can reduce hallucination, but it can also produce a beautifully sourced wrong answer. The source trail matters because it lets the clinician do the checking, not because it removes the need to check.

If you want the broader taxonomy, I put these products into a more complete market map here: The AI in Healthcare Map. For clinical evidence tools specifically, I wrote more about OpenEvidence vs ChatGPT and what OpenEvidence actually does well. And for the more speculative end of the field, see Google's AMIE.

The question is no longer whether AI can produce impressive clinical outputs. It can.

The question is whether the output improves the decision that reaches the patient.

That is still the hard part. Medicine remains annoyingly attached to reality.

Educational content for clinicians. Not medical advice. I have no financial relationship with the companies discussed here. Views are my own.

References

  1. ChatGPT connects health records and healthcare sources. OpenAI. 2026.
  2. 1,357 AI medical devices cleared, 3 actually tested on patient outcomes. PLOS Digital Health. 2026. doi:10.1371/journal.pdig.0001597.
  3. Artificial Intelligence-Enabled Medical Devices. U.S. Food and Drug Administration.
  4. Cost-effectiveness of an AI-based app compared to usual care for early skin cancer detection in Belgium. npj Digital Medicine. 2026. doi:10.1038/s41746-026-03126-y.
  5. Artificial Intelligence-Enabled Acquisition and Interpretation for Screening Aortic Stenosis. JAMA Cardiology. 2026. doi:10.1001/jamacardio.2026.3829.
  6. Vara Receives World-First CE Certification for Autonomous AI in Breast Cancer Screening. Business Wire. 2026.
  7. Nationwide real-world implementation of AI for cancer detection in population-based mammography screening. Nature Medicine. 2025. doi:10.1038/s41591-024-03408-6.
  8. Accuracy of Retrieval-Augmented Large Language Model-Generated Preconsult Summaries in Breast Surgical Oncology. JCO Clinical Cancer Informatics. 2026. doi:10.1200/CCI-26-00160.
Mendel Jacobs, MD, MPH

Mendel Jacobs, MD, MPH

Menachem "Mendel" Jacobs, MD, MPH is an Internal Medicine Resident at Yale School of Medicine pursuing academic cardiology. He publishes under Menachem Jacobs.

Connect on LinkedIn

Related Articles