AI in Medicine

How to Evaluate a Clinical AI Tool: The Five Layers

Medical AI benchmarks contradict each other because they measure different things. A five-layer framework for evaluating clinical LLMs — from board-exam accuracy to patient outcomes — and why the evaluation layer usually decides the winner.

By Mendel Jacobs, MD·Sep 12, 2026·13 min read
Five layers of clinical AI evaluation, from knowledge benchmarks to patient outcomes

The short answer. Studies comparing medical AI tools contradict each other because they measure different things. There are five distinct evaluation layers, and most tools that "win" do so at one layer while losing at another. Some evaluated frontier models led on knowledge and communication tasks; clinical tools performed better on selected citation and harm assessments. Those results depend on the tested versions and methods. Benchmark performance alone does not establish patient benefit. [21] [12]

Disclosure: I use OpenEvidence regularly, including in researching my own writing. No financial relationship with any company mentioned. As of September 2026.


The question people ask isn't the question that matters

"Which is better, OpenEvidence or ChatGPT?" collapses several distinctions into one verdict. It's why the literature offers conflicting answers and why every vendor can find a result worth championing.

A board-style multiple-choice benchmark, a citation-faithfulness audit, a harm-focused consultation benchmark, and a randomized trial of physician use are measuring fundamentally different things. They can point in opposite directions while each remains valid.

The stakes aren't theoretical: 34% of US adults now report using AI chatbots for a health purpose, including 25% to help work out the cause of symptoms [1]. Among physicians, Doximity’s surveyed cohorts reported AI use of 63% in November 2025–January 2026 versus 47% in March–April 2025 [2].

So the useful question is methodological. What are we measuring, and does it predict clinical value?


The five layers

LayerThe questionTypical method
1. KnowledgeDoes it know the answer?MedQA, MMLU, board-style MCQ
2. CommunicationCan it reason and communicate in realistic situations?HealthBench, multi-turn rubrics
3. SafetyCould its recommendation harm someone?Harm-graded consultation benchmarks
4. Clinician performanceDoes a physician decide better using it?Randomized trials of physician use
5. Patient outcomesDoes deploying it change what happens to patients?Pragmatic clinical trials

The layers ask different questions. A result at one layer cannot substitute for evidence at another. Patient-outcome studies are particularly important when the claim concerns benefits in actual care.


Layer 1: interpreting board-exam scores

In one 2026 comparison, the three frontier models scored 90.2–97.4% on a 500-question MedQA sample. Those results are useful, but incomplete as evidence of clinical readiness. [21]

Saturation. High scores can leave less room to distinguish systems, and the aggregate hides which questions they miss. It does not make the benchmark entirely uninformative.

Contamination. One audit detected 99.21% of the MedQA test split inside an open pretraining corpus by 8-gram overlap [3]. A dedicated study found 10–20% verbatim memorization during continued pretraining, with much of it surviving fine-tuning [4]. The counter-evidence is real but partial — a temporal audit found near-identical accuracy on cases published before versus after training cutoffs [5]. Neither overlap in one corpus nor a temporal audit establishes the training history of every model. Contamination needs model- and dataset-specific assessment.

Format inflation. Replace the correct MedQA option with "None of the other answers" and accuracy fell across all six tested models [6]. This demonstrates sensitivity to answer format; it does not by itself prove an absence of reasoning. Deliver the same content conversationally instead of as a tidy vignette and diagnostic accuracy falls again [7].

The knowledge–practice gap, quantified. A systematic review of 39 benchmarks found exam-style accuracy of 84–90% but practice-based success of only 45–69%, with safety performance at 40–50%. These heterogeneous benchmark summaries are not a pooled real-world success rate [8]. On 2,400 real MIMIC-IV cases, models performed significantly worse than clinicians, failed to follow guidelines, couldn't reliably interpret labs even with reference ranges, and were sensitive to prompt changes [9].


Layer 2: communication and information gathering

CRAFT-MD evaluates simulated doctor–patient conversations, including history-taking and diagnostic questioning. Changing a vignette into a dialogue changes the task: the model must obtain information rather than simply receive a complete case. A strong exam score does not measure that ability. [7]

Ask whether a communication benchmark uses real or simulated conversations, who grades the responses, and whether the rubric measures missing information as well as presentation. A clear answer and an adequately supported answer are separate requirements.


Layer 3a: fabrication is not the same as faithfulness

For any retrieval-grounded tool, two failure modes need separating — and conflating them is why one product can be simultaneously praised and criticized for its citations.

Source fabrication: the citation doesn't exist. Common in models without retrieval — one 636-citation study found 55% fabricated for GPT-3.5 versus 18% for GPT-4, with substantive errors also present in 43% of the real GPT-3.5 citations and 24% of the real GPT-4 citations [10].

Faithfulness failure: the citation is real and relevant but doesn't support the specific claim. This is the more insidious mode, because the source trail looks intact. In SourceCheckup, approximately 30% of individual GPT-4o-with-web-search statements were unsupported; about 55% of complete responses were fully supported [11].

Retrieval can improve source access. It does not guarantee faithful synthesis.

That single distinction explains most of the apparent contradiction in vendor claims. "Zero fabricated citations" and "misreads the papers it cites" are both true, and they live on different metrics.


Layer 3b: evaluate safety separately from accuracy

The most important conceptual advance of the past two years is treating safety as its own construct rather than assuming accuracy predicts it.

NOHARM evaluates potentially harmful recommendations and omissions in simulated clinical consultations. Its July 2026 version is a preprint, not a patient-outcome trial. Clinical tools performed better than generalist models on the tested safety measures, but that is not a universal safety certification. [12]

More than 80% of severe errors in that version were omissions. An answer can sound reasonable while leaving out an important action. Earlier versions reported different percentages; results should be cited with their version.

The practical lesson is to assess missing actions as well as incorrect statements.


Layer 4: the clinician-plus-tool problem

Benchmarking the model alone isn't how medicine works. The relevant unit is clinician + tool + workflow, and the evidence here is thinner and stranger.

Access alone did not improve diagnosis in one trial. In the first major RCT, giving physicians GPT-4 plus conventional resources produced no significant improvement in diagnostic reasoning — 76% versus 74% — while the LLM alone beat both physician arms by 16 points [13].

Management reasoning improves modestly. A follow-up trial found a 6.5-point gain on management tasks, though physicians spent more time per case [14].

Workflow design matters. Structured collaborative workflows raised accuracy to 82–85% versus 75%, principally by compressing the low-scoring tail rather than lifting every case [15].

Automation bias survives training. After 20 hours of AI-literacy training, physicians shown deliberately erroneous LLM advice had an adjusted diagnostic-reasoning score 14 percentage points lower — compared with physicians receiving error-free suggestions in simulated cases. This trial tests susceptibility to deliberately incorrect advice, not the overall benefit of routine chatbot access [16].

The collaboration paradox. Across 106 human-AI experiments, when AI alone beats humans, human-AI teams underperform the AI alone. On average, teaming helped more when humans started ahead; this cross-domain meta-analysis does not establish a rule for every clinical task [17].


Layer 5: measure patient outcomes directly

Does deploying the system improve what happens to patients? That question requires evidence from the specific deployment, not an extrapolation from exam performance.

The clearest example now is a pragmatic cluster-randomized trial in Kenya: clinical officers across 16 primary-care facilities randomized to an EMR with or without an LLM copilot, enrolling nearly 10,000 patients. The primary outcome was treatment failure within 14 days — not an exam score, not a physician rating.

Treatment failure: 2.2% with AI versus 2.0% usual care; adjusted odds ratio 0.77 (95% CI 0.55–1.08), P = 0.13. [18].

A US deployment trial, reported as a preprint, randomized 284 outpatient clinicians to a generative chart-summarization tool. It modestly reduced task load and some burnout measures, didn't meaningfully reduce charting time, and clinician engagement declined substantially over 90 days [19].

These examples measure different endpoints. A null result in one setting does not prove that every AI deployment is ineffective, and clinician workload is not a patient health outcome.


What this framework does to the OpenEvidence question

Apply the layers and the apparent contradictions resolve cleanly. The winner depends entirely on which layer you measure.

EvaluationWhat the cited evidence supportsLimit
Knowledge and communicationFrontier models led in one head-to-head studyVersion-, sample-, and method-specific
Citation supportA real source can still fail to support an answerLink validity is not faithfulness
SafetyClinical tools led on the cited NOHARM measuresPreprint; simulated consultation outcomes
Clinician performanceResults differ by task and workflowVignette performance is not patient benefit
Patient outcomesKenya trial found no significant difference in treatment failureOne system, setting, and follow-up period

The head-to-head study also has a published methodological critique and an author response. These discuss contamination, evaluator effects, and the scope of the clinical-query assessment. [21] [22] [23]

Another important variable is the retrieval corpus. A guideline-anchored retrieval configuration beat both an unaugmented frontier model and a literature-only configuration of a dedicated clinical tool on the same cases [20]. This small, specialty-specific study does not establish that the same ordering holds across all tools or specialties.

What a tool can retrieve is part of the intervention being evaluated.

For the practical version of this comparison — which to open at 2 AM — see OpenEvidence vs ChatGPT: What the Evidence Actually Shows.


Six things worth carrying

  1. Match the claim to the layer. "Best medical AI" is meaningless without the metric.
  2. Grounding is not faithfulness. A real citation can still be misread. Demand statement-level support, not a working link.
  3. Inspect the corpus. Retrieval access can affect the result; it is one component of the system.
  4. Safety isn't accuracy, and omissions dominate. The characteristic failure is the missing recommendation.
  5. Deployment is its own experiment. Trial results differ across tasks and workflows; training does not eliminate automation bias.
  6. For the individual clinician: use a grounded clinical tool to reach the evidence fast, then open and read the source before acting. Use frontier chatbots to reason and stress-test — never as a citable source.

The right question is no longer whether AI can answer medical questions. It can. It's whether a given tool improves the specific decision that reaches a specific patient in a messy workflow.

That is the standard patient-facing claims should be held to.


Frequently asked questions

Which medical AI is most accurate? Depends on the metric. Frontier models lead knowledge and communication benchmarks; purpose-built clinical tools lead citation integrity and harm avoidance. No tool leads across all five layers.

Do AI tools improve patient outcomes? The Kenya trial reviewed here found no significant reduction in 14-day treatment failure. That result should not be generalized to every clinical AI application [18].

Why do studies of medical AI contradict each other? Because they measure different things. A board-exam benchmark, a citation audit, a harm benchmark, and a physician RCT can all be valid and point in opposite directions.

Is a high MedQA score meaningful? It measures performance on that exam-style task. Contamination risk and format sensitivity limit how far the result can be generalized [3][6].

What's the difference between hallucination and faithfulness failure? A hallucinated citation doesn't exist. A faithfulness failure is a real citation attached to a claim the paper doesn't support — harder to catch, because the source trail looks intact [11].

Related reading

Educational content for clinicians. Not medical advice. This literature changes monthly; figures as of September 2026.

References

Sources checked September 17, 2026. Preprints are labeled; their findings have not completed peer review.

1. Pew Research Center. From diagnoses to treatments: why Americans use AI chatbots for health. August 25, 2026.

2. Doximity. 2026 State of AI in Medicine Report. Survey report.

3. Gallifant J, et al. Language Models are Surprisingly Fragile to Drug Names in Biomedical Benchmarks. Findings of EMNLP. 2024. doi:10.18653/v1/2024.findings-emnlp.726.

4. Li A, et al. Memorization in large language models in medicine prevalence characteristics and implications. Nature Communications. 2026;17:7729. doi:10.1038/s41467-026-73779-6.

5. Sheppert AP, Adams B, Sheppert AD, Riley S. Reasoning or reciting? A temporal contamination audit of large language models in clinical medicine. JAMIA. 2026. doi:10.1093/jamia/ocag069. PMID:42281378.

6. Bedi S, et al. Fidelity of Medical Reasoning in Large Language Models. JAMA Network Open. 2025;8(8):e2526021. doi:10.1001/jamanetworkopen.2025.26021.

7. Johri S, et al. An evaluation framework for clinical use of large language models in patient interaction tasks. Nature Medicine. 2025. doi:10.1038/s41591-024-03328-5.

8. Gong EJ, Bang CS, Lee JJ, Baik GH. Knowledge-Practice Performance Gap in Clinical Large Language Models: Systematic Review of 39 Benchmarks. Journal of Medical Internet Research. 2025;27:e84120. doi:10.2196/84120.

9. Hager P, et al. Evaluation and mitigation of the limitations of large language models in clinical decision-making. Nature Medicine. 2024;30:2613–2622. doi:10.1038/s41591-024-03097-1.

10. Walters WH, Wilder EI. Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports. 2023;13:14045. doi:10.1038/s41598-023-41032-5.

11. Wu K, et al. An automated framework for assessing how well LLMs cite relevant medical references. Nature Communications. 2025. doi:10.1038/s41467-025-58551-6.

12. Wu D, et al. First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations. arXiv:2512.01241v4. July 13, 2026. PREPRINT, not peer reviewed. doi:10.48550/arXiv.2512.01241.

13. Goh E, et al. Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Network Open. 2024;7:e2440969. doi:10.1001/jamanetworkopen.2024.40969. PMID:39466245.

14. Goh E, et al. GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial. Nature Medicine. 2025;31:1233–1238. doi:10.1038/s41591-024-03456-y.

15. Everett SS, et al. From tool to teammate in a randomized controlled trial of clinician-AI collaborative workflows for diagnosis. npj Digital Medicine. 2026. doi:10.1038/s41746-026-02545-1. PMID:41851268.

16. Qazi IA, et al. Automation Bias in Large Language Model–Assisted Diagnostic Reasoning among Physicians Trained in AI Literacy — A Randomized Clinical Trial. NEJM AI. 2026;3. doi:10.1056/AIoa2501001. Accessible study preprint.

17. Vaccaro M, Almaatouq A, Malone T. When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour. 2024;8:2293–2303. doi:10.1038/s41562-024-02024-1.

18. Agweyu A, et al. Generative AI-enabled clinical decision support system in primary care: a pragmatic, cluster-randomized trial. Nature Medicine. 2026;32:3032–3039. doi:10.1038/s41591-026-04503-6. PMID:42362867.

19. Chin AT, et al. A Pragmatic Randomized Trial of an EHR-Integrated Generative AI Chart Summarization Tool for Ambulatory Clinicians. medRxiv. 2026. PREPRINT, not peer reviewed. doi:10.64898/2026.08.26.26361496.

20. Dukes D, et al. Guideline-anchored retrieval-augmented generation outperforms baseline and literature-only configurations in gynecologic oncology decision support: A pre-integration benchmark. Gynecologic Oncology. 2026;211:224–230. doi:10.1016/j.ygyno.2026.07.005. PMID:42462288.

21. Vishwanath K, et al. General-purpose large language models outperform specialized clinical AI tools on medical benchmarks. Nature Medicine. 2026;32:2405–2409. doi:10.1038/s41591-026-04431-5.

22. Beaulieu-Jones B, Nemati S. Limited benchmarks constrain the conclusions of a general-purpose versus clinical AI comparison. Nature Medicine. 2026. Matters Arising. doi:10.1038/s41591-026-04638-6.

23. Vishwanath K, Aphinyanaphongs Y, Oermann EK. Reply to: Limited benchmarks constrain the conclusions of a general-purpose versus clinical AI comparison. Nature Medicine. 2026. doi:10.1038/s41591-026-04637-7.

Mendel Jacobs, MD, MPH

Mendel Jacobs, MD, MPH

Menachem "Mendel" Jacobs, MD, MPH is an Internal Medicine Resident at Yale School of Medicine pursuing academic cardiology. He publishes under Menachem Jacobs.

Connect on LinkedIn

Related Articles