AI in Medicine

Google's AMIE: What the Evidence Actually Shows

Google's AMIE can now conduct medical conversations, assist cardiologists, and interpret video visits. Here is what the evidence shows, and what it does not.

By Mendel Jacobs, MD·Aug 20, 2026·10 min read
Illustration of AMIE as an AI doctor in a video visit with a patient

Google's medical AI can now see, hear, reason, and conduct a video visit. That is a bigger deal than another chatbot benchmark, but it is still not an autonomous doctor. Maybe a second-year medical student with bad social skills.

My feed has been blowing up with papers and posts about AMIE, Google's Articulate Medical Intelligence Explorer.

And I understand why. For essentially as long as we have had the role of physician, the basic paradigm has been recognizable: a doctor sees a patient, a doctor evaluates a patient, a doctor treats a patient. But the clinical encounter extends far beyond the words exchanged. A physician watches how someone walks into the room. They notice a subtle grimace when the patient moves their shoulder. They register respiratory effort before putting on a stethoscope. They hear whether speech sounds pressured, weak, confused, or breathless. They guide a patient through examination maneuvers and integrate all of this with years of training, pattern recognition, and clinical context.

That sensory layer of medicine has always been one of the strongest arguments against the idea that a chatbot could simply become a doctor. Over the past several years, Google, and more broadly companies developing systems from OpenAI to Anthropic, have steadily chipped away at that limitation. Now Google has taken another substantial step.

On August 10, 2026, researchers released a new preprint describing AMIE Video, a real-time medical AI capable of simultaneously talking with a patient, reasoning about the case, and interpreting audio and video during a telehealth encounter.[1] That release arrives at an interesting moment. A few days later, Ezekiel Emanuel, Vinod Khosla, and colleagues published a provocative JAMA Perspective explicitly asking whether autonomous AI might eventually provide better care than physicians using AI.[2] So are we there? Not remotely.

But AMIE is becoming important because Google is building something resembling an evidence roadmap toward increasingly autonomous clinical interaction, and the progression is worth understanding. It also fits into the broader question I have been tracking in the AI in Medicine pillar: where does clinical AI actually sit in the workflow, what action does it influence, and who stays accountable when it is wrong?

How we got here

2025: Can an AI actually conduct the medical interview?

The first major AMIE milestone came with a Nature study published in April 2025. Researchers compared AMIE with 20 primary care physicians across 159 simulated clinical scenarios using patient actors. Both the physicians and AMIE conducted consultations through text chat. AMIE performed remarkably well. Specialist evaluators rated it superior to physicians across 30 of 32 evaluated domains, and patient actors preferred it across 25 of 26 measures. Diagnostic accuracy was also higher.[3]

Those results generated understandable excitement. They also had an enormous caveat. Doctors do not normally practice medicine through anonymous synchronous text chat.

The investigators themselves emphasized that the study should not be interpreted as representative of normal telemedicine practice. The experiment was arguably better understood as showing that AMIE was extraordinarily good at a particular form of diagnostic conversation, rather than proving it was a better doctor. Still, one barrier had fallen: sophisticated clinical interviewing was no longer obviously a uniquely human capability.

February 2026: What happens when doctors use AMIE?

The next study may actually be more clinically important. In Nature Medicine, researchers evaluated AMIE as a co-pilot for general cardiologists reviewing 107 complex patients with suspected genetic cardiovascular disease. Nine cardiologists reviewed ECGs, echocardiograms, cardiac MRI, cardiopulmonary exercise testing, and other clinical information either with or without AMIE assistance. Blinded subspecialists preferred the AMIE-assisted assessments in 46.7% of cases, compared with 32.7% for cardiologists alone, with the remainder tied. More importantly, clinically significant errors fell from 24.3% to 13.1%, while missing clinically relevant content fell from 37.4% to 17.8%.[4]

That is a meaningful finding. But notice what the experiment actually demonstrated. It did not show that AMIE was a better cardiologist. It showed that a cardiologist with AMIE could outperform a cardiologist without it when reviewing complex retrospective cases. That distinction is enormously important. If these findings reproduce prospectively, one of the first major effects of medical AI may not be replacing specialists. It may be compressing the expertise gradient between generalists and subspecialists.

A community cardiologist faced with an unusual cardiomyopathy could effectively have an additional reasoning layer sitting beside them. That is potentially transformative and life-saving.

2026: Now put AMIE in front of actual patients

Google then moved outside the simulated OSCE. In collaboration with Beth Israel Deaconess Medical Center, researchers prospectively enrolled 100 urgent-care patients who interacted with AMIE through text before seeing their physician. Every AI conversation was monitored by a physician who could stop the encounter for predefined safety concerns.

There were zero formal safety stops. Among the 98 evaluable cases, AMIE's differential included the eventual diagnosis within its first seven possibilities in roughly 90% of cases. Blinded reviewers found no significant difference between AMIE and physicians in the overall quality of the differential, appropriateness of management, or safety. But physicians clearly beat AMIE in two areas that matter enormously in actual medicine: practicality and cost-effectiveness.[5]

That result should probably receive more attention than it has. Knowing the medically possible answer is not the same as knowing what should actually happen to this particular patient, in this health system, with this insurance plan, after everything that has already happened to them.

AMIE did not have full access to that context. Medicine is partly diagnosis. It is also logistics. It is also getting a vibe from your patient to understand how and if your "correct" medical advice suits their needs in that moment.

And now AMIE can see

That brings us to the newest experiment. Google's August 2026 preprint moves AMIE from text into real-time audio-visual telemedicine. The study included 100 clinical scenarios, enacted by 15 professional patient actors. Ten board-certified primary care physicians conducted video consultations, while another 20 PCPs independently evaluated the encounters. AMIE Video was compared both with physicians and with AMIE's own text-only configuration.[1]

A simplified PICO looks like this:

  • Population: 100 simulated clinical scenarios containing audio-visual findings and virtual examination components.
  • Intervention: AMIE Video, a Gemini-based AI conducting synchronous audio-video consultations.
  • Comparator: Board-certified PCPs conducting video consultations, plus text-only AMIE.
  • Outcomes: Clinical history-taking, diagnosis, management, physical observation, virtual examination, communication, and patient-actor preferences.

The headline result is impressive. Independent clinician evaluators rated AMIE Video on par with or better than physicians across history taking, diagnosis, management, physical observation, and physical examination. Patient actors preferred AMIE's approach to assessing and explaining their conditions, although physicians retained an advantage in rapport and partnership building.[1]

The most interesting part is not actually the medical score

To me, the most important innovation may be the architecture. A major problem with putting sophisticated reasoning models into live conversation is latency. Clinical reasoning takes computation. Conversation cannot tolerate someone staring silently into space for 20 seconds every time you answer a question. Google's solution is essentially to split the AI physician into several processes working at once.

There is a Talker, optimized for rapid conversation.

A Planner works in the background, maintaining the differential diagnosis, management plan, unanswered questions, and goals of the encounter.

And a Perception agent continuously evaluates the audio-visual stream for clinically relevant cues: affect, speech cadence, cough, movement, visible findings. Those systems communicate asynchronously. In other words, the AI can continue "thinking" while another part of it continues talking. That begins to look less like ChatGPT with a webcam and more like an actual clinical-agent architecture.[1]

This is impressive. It is also still an OSCE.

What the evidence does not show yet

This is where some of the commentary online gets ahead of the evidence. The patients in the video trial were professional actors. That is not a trivial limitation. Actors follow scripts. Real patients forget medications, contradict themselves, bring in their spouse, misunderstand questions, have hearing impairment, speak through interpreters, have five diseases simultaneously, become frustrated, cry, wander off camera, or tell you that the chest pain started yesterday before later remembering it began last week.

Clinical medicine is extraordinarily adversarial data. A well-designed OSCE tests clinical skill and is meant for third-year medical students before their school trusts them taking to a real human. It does not recreate the full entropy of an actual clinic.

There is another important distinction: these studies largely measure intermediate clinical performance, not health outcomes. No one has randomized thousands of patients with chest pain to AMIE versus physicians and demonstrated fewer missed myocardial infarctions, fewer unnecessary emergency referrals, lower costs, or improved mortality.

The appropriate evidence ladder looks something like this:

benchmark -> simulation -> clinician assistance -> supervised real-world deployment -> prospective comparative trials -> patient outcomes -> autonomous care.

AMIE has climbed considerably higher on that ladder than most medical LLMs. But it has not reached the top. There are also things the system still cannot do well. The paper itself identifies persistent weaknesses in recognizing fine anatomical details, subtle affective signals, and rapid movements.[1]

And a screen remains a screen. An AI can ask someone to demonstrate shoulder range of motion or pronator drift. It cannot actually feel abdominal rigidity. It cannot palpate a pulse. It cannot appreciate the texture of a rash beyond what the camera captures.

It cannot yet reproduce the enormous stream of tactile, spatial, contextual, and social information contained in an in-person encounter. There is also the larger problem of contextual intelligence.

The BIDMC study is instructive here. AMIE could generate clinically plausible recommendations while still producing plans that reviewers considered less practical and less cost-conscious than physicians' plans.[5] A truly autonomous clinical agent will eventually need much more than medical knowledge.

So, will AI become a better doctor than doctors?

That is the question Emanuel and colleagues are now willing to ask openly in JAMA. It is no longer an absurd question.[2] But it is also not a question these AMIE studies answer. What Google has demonstrated is narrower and, in some ways, more interesting.

AI can conduct sophisticated diagnostic dialogue. AI can improve physician reasoning in difficult specialty cases. AI can interact safely enough with real patients to justify carefully supervised prospective studies. And AI can now combine clinical reasoning with real-time sight, sound, and conversation at a level that approaches physicians in a controlled telehealth simulation.

Those are substantial achievements. They are not equivalent to demonstrating autonomous medical practice. For the immediate future, the more plausible model is probably neither doctor versus AI nor even merely doctor plus chatbot.

It is an increasingly integrated clinical system in which AI listens to the encounter, sees what is happening, retrieves the record, builds a differential in parallel, notices omissions, suggests examination maneuvers, checks guideline-based care, and quietly warns the physician when something does not fit.

That could change medicine considerably before a single physician is "replaced." And that, to me, is what makes AMIE worth watching. The remarkable thing is not that Google has created an AI doctor. It has not.

The remarkable thing is how quickly the list of things we once assumed only a doctor could do is getting shorter.

For a broader taxonomy of where tools like AMIE fit, see The AI in Healthcare Map.

References

  1. Towards Conversational Diagnostic AI in Real-World Multimodal Clinical Settings. arXiv. 2026.
  2. Autonomous AI and the Future of Medical Care. JAMA. 2026.
  3. AMIE: a research AI system for diagnostic medical reasoning and conversations. Nature. 2025.
  4. AMIE-assisted specialist assessment for suspected genetic cardiovascular disease. Nature Medicine. 2026.
  5. Exploring the feasibility of conversational diagnostic AI in a real-world clinical study. Google Research Blog. 2026.
Mendel Jacobs, MD, MPH

Mendel Jacobs, MD, MPH

Menachem "Mendel" Jacobs, MD, MPH is an Internal Medicine Resident at Yale School of Medicine pursuing academic cardiology. He publishes under Menachem Jacobs.

Connect on LinkedIn

Related Articles