AI in Medicine

OpenEvidence vs ChatGPT: Which Should a Clinician Actually Use?

OpenEvidence versus ChatGPT: compare evidence on medical benchmarks, citation reliability, safety, and physician use, with study-specific limitations.

By Mendel Jacobs, MD·Sep 12, 2026·10 min read
OpenEvidence versus ChatGPT for clinical use: what the evidence supports

The short answer. There is no single winner across every task. In the studies reviewed here, tested frontier models performed better on selected knowledge and communication measures, while clinical tools showed advantages on particular citation and safety assessments. Those findings apply to the versions and tasks evaluated. They do not establish that either product is appropriate for every clinical decision.

Disclosure: I use OpenEvidence regularly, including in researching my own writing. No financial relationship with any company mentioned. As of September 2026.


Why the studies seem to disagree

They don't, really. They measure different things.

A board-style benchmark, a citation audit, a harm-focused consultation benchmark, and a randomized trial of physician use are four different experiments. Each can be valid while pointing in a different direction — and each vendor can truthfully cite the one that favors them.

I've laid out the five layers of clinical AI evaluation separately. The short version:

LayerWhat it asks
1. KnowledgeDoes it know the answer?
2. CommunicationCan it reason and communicate realistically?
3. SafetyCould its recommendation harm someone?
4. Clinician performanceDo you decide better using it?
5. Patient outcomesDoes deploying it change what happens?

Apply those and the contradiction dissolves. The winner changes by layer.


Layers 1–2: frontier models led in one head-to-head comparison

The most direct head-to-head compared OpenEvidence and UpToDate's Expert AI against GPT-5.2, Gemini 3.1 Pro Preview, and Claude Opus 4.6. The study sampled 500 MedQA and 500 HealthBench items. [1]

  • MedQA: Gemini 97.4% vs OpenEvidence 89.6%, UpToDate 88.4%
  • HealthBench: GPT 88.0 vs OpenEvidence 62.6, UpToDate 61.3 — clinical tools last in all seven thematic categories
  • 100 real de-identified physician queries, scored by 12 blinded clinicians: frontier models 3.5–3.6/4, clinical tools both 3.2

Two details in that last one deserve attention. UpToDate's tool refused 19% of queries. And Google's Search AI Overview scored 3.3 — matching the dedicated clinical tools.

Binary harm and hallucination flags did not differ significantly across systems.

The authors themselves caution this is a snapshot, that citation quality and latency weren't assessed — dimensions that likely favor purpose-built tools — and that HealthBench's provenance raises developer-overlap concerns, since it was built by one of the competitors.

A September 2026 critique raised concerns about benchmark contamination, evaluator affinity, and the clinical-query evaluation. The authors replied that their main conclusion holds under the tested conditions. Read both alongside the original paper. [2] [3]


Layer 3a: citation reliability is a separate outcome

Where verifiability is the metric, the picture inverts.

In a cervical-spine guideline evaluation, OpenEvidence produced 999 citations with 100% accuracy, while only 46.8% of ChatGPT-4o's citations met accuracy criteria. The study tested OpenEvidence 2.0 and ChatGPT-4o on cervical-spine guideline questions, not every kind of clinical query. [4]

Across otolaryngology vignettes, ChatGPT-4 was numerically the most diagnostically accurate (91% vs 82–89%, not significant) — and had a reported citation hallucination rate of 23%, while OpenEvidence provided more consistent references and had the highest mean CiteScore for journal citations. CiteScore is a journal-level metric, not a guarantee that an individual citation supports a claim. [5]

That's the consistent trade-off: frontier models may edge out diagnostic accuracy while clinical retrieval tools deliver sources you can check.

But there's a second failure mode that matters more. A fabricated citation is obvious once you click it. A real citation attached to a claim the paper doesn't support is not — and SourceCheckup found that about 30% of individual statements from GPT-4o with web search were unsupported. Roughly 55% of complete responses were fully supported. Those are different denominators. [6]

Retrieval can provide a route to the evidence, but it does not guarantee that the answer accurately represents it.

So: open the citation. Then read enough of it to confirm it says what the tool claims.


Layer 3b: clinical tools cluster at the top on harm

The July 2026 NOHARM preprint found lower potential harm for the clinical tools it evaluated than for the generalist models. More than 80% of severe errors were omissions. These are simulated consultation assessments, not observed patient injuries or a universal product safety ranking. [7]

The practical implication is to check what an answer leaves out as carefully as what it recommends.


The retrieval corpus is part of the comparison

Product name alone does not describe the system: what it can retrieve also matters.

On 50 gynecologic oncology cases, a guideline-anchored configuration scored 0.83, significantly beating both a baseline frontier model (0.65) and a literature-only configuration of a dedicated clinical tool (0.70) that lacked guideline access at test time. The clinical tool and the baseline model didn't differ significantly from each other. [8]

The comparison preceded guideline integration in the clinical tool. It should not be treated as a current product ranking or proof that the same retrieval hierarchy holds across specialties.


Layer 4: what happens when you actually use it

Benchmarking the model alone isn't how medicine works.

Access alone did not improve diagnosis in one trial. Physicians given GPT-4 plus conventional resources showed no significant improvement in diagnostic reasoning — 76% vs 74% — while the model alone beat both physician arms by 16 points. [9]

Management reasoning improves modestly. A follow-up trial found a 6.5-point gain on management tasks, at the cost of more time per case. [10]

Automation bias survives training. After 20 hours of AI-literacy training, physicians shown deliberately erroneous advice had an adjusted diagnostic-reasoning score 14 percentage points lower. The comparison was with error-free suggestions in a simulated task; it does not measure the net benefit of ordinary chatbot access. [11]

And the collaboration paradox: across 106 human-AI experiments, when AI alone beats humans, human-AI teams underperform the AI alone. That pattern was an average across diverse tasks, not a rule for every clinical setting. [12]


Layer 5: patient benefit requires a different trial

The comparative benchmark studies above do not establish patient-outcome benefit for either product.

The clearest attempt — a pragmatic cluster-randomized trial across 16 primary-care facilities, nearly 10,000 patients — found treatment failure at 14 days of 2.2% with an LLM copilot versus 2.0% usual care. The adjusted estimate was not statistically significant, and its confidence interval included both benefit and harm. [13]

That sentence should calibrate every claim above, including the favorable ones.


What I'd actually tell a resident

Start with evidence you can inspect. A clinical search tool may help locate sources, but treatment decisions require checking the evidence, patient context, and appropriate clinical guidance.

Check the current question and context. Make sure a retrieved answer addresses the actual population and decision in front of you. Do not assume that a citation attached to an earlier answer supports a later one.

Use frontier models to think, not to cite. Differential building, "what am I missing," explaining a mechanism to yourself — these are possible uses to evaluate, not tasks for which every frontier model has proven superiority. Never treat their output as a source.

Never trust either one's summary of a paper you haven't opened. The documented failure mode of the best tool is misreading its own citations.

Watch yourself, not just the tool. The omission finding and the automation-bias data both point the same way: the risk isn't that you'll be told something absurd. It's that you'll accept something incomplete because it reads well.

If a patient asks: explain the difference between finding general information and making a personal medical decision. An apparently well-sourced answer still needs clinical context.


What would change my mind

A prospective trial showing that physician-plus-tool decisions are better for patients than physician alone — on outcomes, not vignettes. If that lands for either category, this article gets rewritten.

A negative deployment trial would also need to be interpreted in its setting, rather than generalized to every use of clinical AI.

Nobody knows which yet. Be suspicious of anyone who claims to.


Frequently asked questions

Is OpenEvidence better than ChatGPT? The answer depends on the task and tested version. The cited comparison favored frontier models on selected benchmarks, while other studies favored clinical tools on citation or harm measures. None of those metrics alone establishes better patient care.

Does OpenEvidence hallucinate? The cited specialty studies found favorable citation reliability for OpenEvidence. That does not mean every answer is correct or that a real citation necessarily supports the surrounding claim.

Is ChatGPT safe for medical questions? The studies here do not establish universal safety. Check the source and clinical context, and do not use a chatbot as a substitute for professional medical assessment.

Which should I use for a clinical question at 2 AM? Use appropriate clinical resources and local guidance. If an AI tool helps find evidence, open the citations and verify that they support the proposed answer.

Do any of these improve patient outcomes? The Kenya trial reviewed here found no significant reduction in 14-day treatment failure. That is evidence about one LLM-supported workflow, not a head-to-head patient-outcome trial of OpenEvidence and ChatGPT.

Related reading

Educational content for clinicians. Not medical advice. This field changes monthly; figures as of September 2026.

References

Sources checked September 17, 2026. Preprints and conference versions are identified below.

1. Vishwanath K, et al. General-purpose large language models outperform specialized clinical AI tools on medical benchmarks. Nature Medicine. 2026;32:2405–2409. doi:10.1038/s41591-026-04431-5.

2. Beaulieu-Jones B, Nemati S. Limited benchmarks constrain the conclusions of a general-purpose versus clinical AI comparison. Nature Medicine. 2026. Matters Arising. doi:10.1038/s41591-026-04638-6.

3. Vishwanath K, Aphinyanaphongs Y, Oermann EK. Reply to: Limited benchmarks constrain the conclusions of a general-purpose versus clinical AI comparison. Nature Medicine. 2026. doi:10.1038/s41591-026-04637-7.

4. Avrumova F, et al. Evolution of Generative Artificial Intelligence in Clinical Practice: Comparative Performance of OpenEvidence 2.0 and ChatGPT-4o. The Spine Journal. Published online August 31, 2026. doi:10.1016/j.spinee.2026.08.002. PMID:42674131.

5. Pennington-FitzGerald W, et al. Diagnostic accuracy and citation integrity of four large language models on otolaryngology vignettes. European Archives of Oto-Rhino-Laryngology. 2026;283:4707–4714. doi:10.1007/s00405-026-10253-5. PMID:42120572.

6. Wu K, et al. An automated framework for assessing how well LLMs cite relevant medical references. Nature Communications. 2025. doi:10.1038/s41467-025-58551-6.

7. Wu D, et al. First, do NOHARM: a medical safety benchmark and randomized study of physician and AI teaming on clinical consultations. arXiv:2512.01241v4. July 13, 2026. PREPRINT, not peer reviewed. doi:10.48550/arXiv.2512.01241.

8. Dukes D, et al. Guideline-anchored retrieval-augmented generation outperforms baseline and literature-only configurations in gynecologic oncology decision support: A pre-integration benchmark. Gynecologic Oncology. 2026;211:224–230. doi:10.1016/j.ygyno.2026.07.005. PMID:42462288.

9. Goh E, et al. Large Language Model Influence on Diagnostic Reasoning: A Randomized Clinical Trial. JAMA Network Open. 2024;7:e2440969. doi:10.1001/jamanetworkopen.2024.40969. PMID:39466245.

10. Goh E, et al. GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial. Nature Medicine. 2025;31:1233–1238. doi:10.1038/s41591-024-03456-y.

11. Qazi IA, et al. Automation Bias in Large Language Model–Assisted Diagnostic Reasoning among Physicians Trained in AI Literacy — A Randomized Clinical Trial. NEJM AI. 2026;3. doi:10.1056/AIoa2501001. Accessible study preprint.

12. Vaccaro M, Almaatouq A, Malone T. When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour. 2024;8:2293–2303. doi:10.1038/s41562-024-02024-1.

13. Agweyu A, et al. Generative AI-enabled clinical decision support system in primary care: a pragmatic, cluster-randomized trial. Nature Medicine. 2026;32:3032–3039. doi:10.1038/s41591-026-04503-6. PMID:42362867.

Mendel Jacobs, MD, MPH

Mendel Jacobs, MD, MPH

Menachem "Mendel" Jacobs, MD, MPH is an Internal Medicine Resident at Yale School of Medicine pursuing academic cardiology. He publishes under Menachem Jacobs.

Connect on LinkedIn

Related Articles