Skip to main content
The Quantum Dispatch
Back to Home
Cover illustration for AMIE Video Consultations Match Doctors in a New Study

AMIE Video Consultations Match Doctors in a New Study

Google's AMIE matched or beat 30 primary care physicians across 100 clinical scenarios in the first real-time video consultation study of its kind.

Dr. Nova Chen
Dr. Nova ChenAug 12, 20266 min read

A Study Design That Actually Tests the Hard Part

Text-based medical AI has been evaluated to death. Video is a different problem, because a consultation over video includes things a chat window cannot carry: a rash you can see, a swelling you can compare side to side, the way someone moves when they stand up. On August 12, 2026, Google published results from what it describes as the first study of expert-level AI in real-time clinical video consultations, using its AMIE research system.

  • 30 primary care physicians, 15 patient actors, and 100 clinical scenarios in a randomized OSCE-style study
  • AMIE (Video) rated on par or better than physicians on history-taking, diagnosis, management, and physical observation
  • Physicians were still preferred for rapport and partnership building
  • Built on Gemini as a multi-agent system combining low-latency dialogue, clinical reasoning, and real-time audio-visual perception

The structure matters as much as the result. An Objective Structured Clinical Examination is the format medical schools use to assess actual clinicians, with trained patient actors running standardized scenarios and independent evaluators scoring the encounter. Running AMIE through the same gauntlet, randomized against real physicians doing the same scenarios over video, is a considerably harder test than a benchmark of exam questions.

What Did the AMIE Video Study Find?

Clinical evaluators rated AMIE (Video) on par with or better than the primary care physicians across four dimensions: taking a history, reaching a diagnosis, proposing management, and physical observation and examination. That last one is the novel dimension — it is the thing the video channel exists to enable, and it is where a text-only system has nothing to work with.

The comparison also included AMIE (Text), the system's chat-only counterpart, which lets the researchers separate "the model got better" from "the video channel helped." Patient actors preferred the video interface over text chat on communicative effectiveness, convenience, and feeling understood.

The result I would underline, though, is the one that went the other way: physicians were preferred for rapport and partnership building. Patient actors liked how AMIE assessed and explained their conditions, and preferred the humans for the relational parts of care. That is a clean, believable split, and it is a more useful finding than a clean sweep would have been.

Why Does Real-Time Video Change the Problem?

Because latency and perception have to work together, continuously, while a conversation is happening.

A text medical AI can take several seconds to reason. A video consultation cannot — a pause that long reads as a dropped call. AMIE is described as a multi-agent system precisely because those jobs get separated: something maintains the dialogue at conversational speed, something else does the slower clinical reasoning, and a perception component processes what the camera is seeing in real time.

That architecture is becoming the standard answer for agentic systems that have to be both fast and careful, and it shows up well outside medicine. It is the same decomposition behind a lot of current agent tooling, where a quick responder sits in front of a deliberate reasoner.

What This Does Not Mean

It is worth being precise, because medical AI headlines rarely are. This is a research system in a study with patient actors, not a product, not a deployment, and not a substitute for seeing a clinician. Standardized scenarios are designed to be assessable, which makes them cleaner than real patients, who arrive with comorbidities, incomplete histories, and questions that do not fit the script.

What the study does establish is that the video channel is worth pursuing — that adding real-time visual perception to a conversational medical AI measurably improves the things you would hope it improves, and that patients find it easier to talk to than a chat box. Given how much of global healthcare access is bottlenecked on clinician time, that is a meaningful direction even at this stage.

It also lands in a week where AI is showing up across the research pipeline rather than just the clinic. Our coverage of AWS and Novo Nordisk putting agents into drug discovery looks at the other end of the same trend, and earlier work like AI surfacing 1,000+ antibiotic candidates from prion proteins shows how quickly the discovery side has moved. More in our AI coverage.

Sources: Google Blog — August 12, 2026; Google Research — 2026; arXiv 2608.09861 — August 2026.

More AI Stories