A Benchmark Cannot Grant Permission
Google calls AMIE expert-level. Its own work shows why capability is not authority.
Google has a new medical AI paper, and the research is strong enough that the framing deserves scrutiny. “Towards Expert-level Medical AI for Real-time Video Consultations” comes from Google Research and Google DeepMind. The title contains meaningful restraints, including “towards” and the limitation to video consultations, while the abstract goes further and describes the work as the first demonstration of expert-level AI in real-time clinical video consultations.
“Expert-level performance” can be a legitimate description of benchmark results when the benchmark remains attached to the claim. Trouble begins when the experimental boundary falls away while “expert-level medical AI” keeps traveling. Real-time clinical video consultation sounds remarkably close to clinical practice. What Google actually tested was narrower, cleaner, and substantially easier to govern.
The underlying system is genuinely impressive. AMIE Video combines agents for dialogue, clinical planning, and real-time audio-visual perception, and across Google’s evaluation it performed extremely well against primary-care physicians on multiple measures. Those results deserve attention. They also require precision about what was demonstrated because medicine already knows something the AI industry keeps having to relearn. Competence and authority are related, but they are not interchangeable.
What the benchmark actually measured
Google evaluated AMIE using 100 Objective Structured Clinical Examination (OSCE) scenarios. An OSCE makes clinical assessment comparable by giving participants standardized cases, controlled conditions, and defined scoring criteria. That structure is exactly what makes the method useful for research, and exactly what separates it from ordinary clinical practice.
Professional patient actors received structured scenario packs containing expected diagnoses, plausible alternatives, investigations, management approaches, and scoring rubrics. The study excluded inpatient cases, pediatric cases, dermatological cases, and cases requiring privacy-sensitive physical examinations. Those exclusions are reasonable for a controlled video-based experiment, but they also define the boundary of the evidence. The study concentrates on problems particularly suited to structured remote interaction and standardized evaluation.
Within that envelope, AMIE’s performance is meaningful. It does not establish general medical expertise, real-world patient outcomes, or autonomous clinical competence across the messier distribution of cases that arrives without a scenario pack. The authors acknowledge that further research is required before real-world translation, so the limitation is not hidden. The asymmetry lies elsewhere. Prestige language is available immediately, while the boundary conditions require the reader to understand the study design.
The phrase “expert-level” is not the problem by itself. The problem is what happens when “on standardized primary-care video scenarios” disappears while “expert-level medical AI” keeps traveling through press coverage, executive decks, investment narratives, and eventually deployment arguments.
The human comparison is narrower than it first appears
The abstract says the study involved 30 primary-care physicians, 15 patient actors, and 100 scenarios. That statement is accurate, but it can land more broadly than the methods support. Ten board-certified primary-care physicians actually conducted the comparison consultations, while the remaining twenty physicians served as evaluators. Thirty physicians participated in the study, but AMIE was not compared against thirty physicians performing the encounters.
This does not invalidate the research, and it should not be treated as a cheap gotcha. It matters because the composition of the human comparison group contributes directly to the meaning readers assign to “expert-level.” When the prestige claim depends partly on comparison with physicians, the structure of that comparison deserves the same prominence as the result.
The encounters carried additional experimental constraints. Physicians were instructed to keep their cameras off, while AMIE had no visual avatar. The patient remained visible to both, but the reciprocal visual interaction normally present in a video consultation was removed. That choice reduces one source of evaluation bias, but every control has a price. In this case, the human encounter becomes somewhat less representative of ordinary telehealth.
After each consultation, both AMIE and the physicians completed structured assessments covering differential diagnosis, observations, and management recommendations. For that stage, Google gave AMIE additional inference-time compute using multi-draft synthesis to improve recommendation quality. The paper does not describe an equivalent augmentation for the physicians. Again, the experiment remains useful. The object simply becomes more precise than the phrase “AI versus doctor” implies.
Google reports that AMIE placed the predetermined reference diagnosis first in 91 percent of scenarios compared with 77 percent for physicians. When the top three differential diagnoses were considered, the comparison narrowed to 98 percent versus 90 percent and was not statistically significant at the reported threshold. Evaluators also rated AMIE highly across history-taking, clinical reasoning, management planning, communication, and guided examination.
Those are substantial results, and they should be treated as such. AMIE may already demonstrate expert-level performance on some standardized clinical tasks. That can be true while telling us remarkably little about how much clinical authority the system should receive. A system can outperform a physician on a bounded task and still warrant less authority because authority depends on more than task performance.
Google already knows the distinction
Google’s prior AMIE work uses a cleaner boundary between what the system can do and what it is allowed to do. In August 2025, Google described a guardrailed version of AMIE for clinical history-taking and explicitly recognized that individualized diagnosis and treatment are regulated activities requiring licensed medical review. The system was constrained from directly providing individualized medical advice and instead generated material for physician review.
That is good control design because the system’s demonstrated capability does not silently expand its permission. In March 2026, Google described prospective real-world AMIE research at Beth Israel Deaconess Medical Center. Again, AMIE was not installed as an autonomous physician. It performed pre-visit history-taking while physicians retained oversight, and Google described translation into practice as requiring a safety-centered and evidence-based process.
The organization plainly understands the distinction. Generating a plausible diagnosis does not confer independent diagnostic authority. Producing a treatment recommendation does not create prescribing authority. Moving from structured evaluation to real patients changes the evidence burden because the consequence surface changes with it.
The new paper does not erase that discipline, but its framing travels farther than those boundaries. “Expert-level” carries reputational weight immediately while authority, liability, deployment, and patient outcomes remain subjects for later evidence.
The capability claim arrives first. The governance bill arrives later.
Medicine already separates competence from authority
AI does not need a special restriction invented because machines make people nervous. Medicine already separates competence from authority for humans, and it does so because skill has never been sufficient grounds for unrestricted action.
A physician does not gain permission to practice anything they are capable of understanding. Clinical authority is bounded through law, licensure, credentialing, institutional privileging, scope of practice, supervision, professional standards, and patient consent. A brilliant cardiologist does not acquire neurosurgical privileges because they can pass a neurosurgery examination. Knowledge can support a grant of authority. It cannot issue the grant.
A benchmark can inform a delegation decision. It cannot perform the delegation. That distinction matters more as AI improves because AI allows us to separate outputs associated with expertise from the professional obligations historically attached to producing them.
A diagnosis from a physician arrives inside an institutional structure. There is a licensed person, a professional duty, a defined scope of practice, a medical record, a liability structure, a patient relationship, and mechanisms for review. Those structures do not make clinicians infallible. They make authority legible enough that responsibility has somewhere to land.
AI can reproduce some outputs of expertise while leaving duty, liability, and obligation outside the model. That is the deeper governance problem. If clinical authority is delegated to an AI system, someone still has to grant it, bound it, monitor it, revoke it when necessary, and answer for what happens when the system acts within that delegation and gets the decision wrong.
The patient belongs in the authority chain
Medical AI discussions have a convenient habit of turning into contests between model scores and physician scores while the person receiving care slowly disappears from view. The patient is not simply the endpoint of a clinical pipeline. They have an interest in knowing who or what materially shaped a consequential decision, whether a human can reconsider it, whether the recommendation can be challenged, what information was used, and who is accountable when the result causes harm.
Patient consent will not solve every governance problem. A checkbox saying AI may be used in care becomes liability theater if the system’s actual role remains illegible. Meaningful consent requires meaningful boundaries. If an AI is taking a history, that can be stated plainly. If it is generating a differential diagnosis for physician review, the role can be named just as clearly. As the system gains authority to initiate orders, modify treatment, prioritize access, or make other consequential choices, the delegation chain should become more explicit rather than less.
Sovereignty for patients requires liability for operators. Otherwise “human oversight” becomes a phrase everyone can invoke and nobody can locate when the decision needs to be challenged.
SAFE asks for receipts
The Shared AI Findings Exchange (SAFE) proposal from the Open Secure AI Alliance comes from security rather than medicine, but its useful contribution is operational. SAFE does not begin by asking whether an agent is intelligent, aligned, impressive, or expert-level. It asks what happened, what the system could access, what boundaries were crossed, and what evidence survived.
The proposal calls for preserving prompts, traces, tool calls, configuration, model versions, agent identities, available permissions, human approval events, timelines, reproduction evidence, and remediation evidence. Recommendations are supposed to identify the failure, the required defensive outcome, a verification method, retained evidence, a responsible owner, and a deadline.
None of that makes logging equivalent to governance. Perfect forensic documentation of a failure nobody had authority to prevent is still failure with excellent paperwork. SAFE matters because it treats evidence as something that must survive consequential action. Medical AI needs the same property. When a system influences care, operators should be able to reconstruct what it perceived, inferred, recommended, what authority it possessed, whether approval was required, who approved it, and what action followed.
Without that chain, “the AI recommended it” becomes an accountability dead end. The conclusion survives while responsibility evaporates.
POLIS asks where permission came from
“Multi-Agent AI Safety as an Institutional Design Problem” approaches the same fault line through authority provenance. Across 5,280 structured experimental episodes, the researchers examined how rules, guards, delegation, and authority interact. In one matched authority-laundering condition, a guard relying on mutable local policy permitted prohibited actions in 22 of 96 episodes, while a provenance-aware guard retaining the originating authority state permitted none.
Those numbers are not predictions about healthcare. POLIS is a structured experimental environment, not a medical deployment study, and treating its failure rates as transferable would commit exactly the benchmark inflation this article is criticizing.
The useful mechanism is narrower. A downstream representation of permission can change without the originating authority changing. If enforcement trusts only the latest local state, a restriction can effectively disappear during delegation. Provenance allows the system to ask where permission came from and whether the actor granting it had the authority to do so.
Infrastructure engineers already know this problem. A service account does not acquire database-admin rights because it became extremely good at SQL, and a Kubernetes workload does not receive cluster-admin because it passed a benchmark. Authorization comes from an authority outside the workload. Medical AI needs the same separation. Competence should influence what we are willing to delegate without ever being confused with the delegation itself.
What would change the judgement
This is not a permanent veto against clinical AI. There is evidence that would justify increasing authority, and being explicit about that threshold matters because otherwise every objection can be dismissed as a gate that will move as soon as the technology improves.
A serious case for wider AMIE deployment would need prospective evidence across real patients and multiple clinical settings, not only standardized encounters. It would need credible performance across demographic and clinical subgroups, behavior under uncertainty and out-of-distribution conditions, reliable escalation beyond its competence envelope, and patient outcomes rather than reference-answer agreement alone.
It would also need an explicit delegation architecture. Operators should be able to answer which actions the system can take independently, which require a physician, who defines those permissions, how they are revoked, what happens when the model and clinician disagree, what evidence is retained, who owns the resulting decision, and where a patient can seek human review. “Human in the loop” is not an accountability model if the human lacks the time, information, or practical authority to disagree.
Those controls have a price. They add engineering work, clinical workload, validation, operational friction, audit burden, and potentially slower deployment. Good. High-consequence authority is supposed to be expensive, and if the economics only work when oversight becomes ceremonial, the economics do not work.
Future multi-site evidence showing strong patient outcomes, subgroup safety, calibrated uncertainty, reliable escalation, auditable delegation, meaningful human intervention, and clear accountability should change the amount of clinical authority justified for systems like AMIE. That is a threshold, not an endlessly moving gate.
The actual achievement
Google’s AMIE research may prove historically important. The system is already capable enough that dismissing it as a chatbot would be ridiculous, and it may eventually outperform human clinicians across substantial portions of diagnostic work. That possibility makes the authority question more urgent because the easiest systems to over-trust are not the incompetent ones. They are the ones that are usually right.
If the public claim were simply “expert-level performance on standardized primary-care video scenarios,” much of my objection would disappear. The description is less dramatic, but it is also closer to what the study establishes. The harder problem begins when performance becomes a basis for permission.
Medicine has spent generations binding knowledge to duty. Its institutions frequently fail, and physicians are hardly immune from bad judgement, but the underlying principle remains sound. The person permitted to act is supposed to be identifiable, their authority is supposed to have boundaries, and their decisions create obligations that someone must carry.
AI makes it technically possible to unbundle the output from the duty. Doing that deliberately may eventually be justified in some clinical contexts, but doing it accidentally because benchmark competence acquired the language of institutional expertise would be governance by semantic drift.
Google does not get to collapse the authority question into the word “expert.” Neither should anyone else. Before an AI system moves from answering like a clinician to acting like one, somebody has to grant that authority, bound it, witness it, preserve the evidence, and accept the consequences.
Capability is not authority. A benchmark cannot grant permission.
Artifacts are cheap, judgement is scarce.
Per ignem, veritas.



