Human-AI clinical teams: choose the right comparison

Before asking whether AI is better than a doctor, ask which combination was tested. A clinician with a chatbot, a model answering a prepared case, and a redesigned care pathway are different interventions. Their results answer different questions.

Better than humans is not the same as better than either alone

A systematic review of 106 experiments distinguished augmentation, where the combination beats humans alone, from synergy, where it beats the better of humans or AI alone. Combinations underperformed the better constituent on average, but results varied across tasks. The experiments spanned multiple domains and were published through June 2023; they are not 106 clinical trials or a verdict on current medical models. Use the distinction to read a claim, not to declare collaboration a success or failure in every setting. [3]

Access to a chatbot is an intervention, not a complete workflow

In a randomized diagnostic-reasoning study of 50 physicians, GPT-4 access alongside conventional resources did not significantly improve scores over conventional resources alone. The task used clinical vignettes, not live patient encounters. The model-alone comparison was exploratory and used a researcher-developed prompt. A high model score therefore does not establish safe autonomous diagnosis, and the physician comparison does not establish that every form of assistance is ineffective. Keep the interface, instructions and task attached to the result rather than treating model capability as a finished clinical service. [1]

A better answer can take more time

A separate randomized study of 92 physicians tested management reasoning across five simulated clinical cases. GPT-4 assistance improved adjusted rubric scores by 6.5 percentage points relative to conventional resources, but physicians spent about 119 seconds longer per case. The assisted group did not significantly outperform the model alone. Those findings concern scored decisions and time in a simulation, not observed patient benefit or a demonstrated staffing saving. In a deployment proposal, ask where the extra work would fall and whether the expected benefit is worth that burden. [2]

Read the whole comparison before choosing a tool

Our reading approach is to write down the actual alternatives: usual care, the proposed assisted workflow, and any model-only arm that was evaluated. Check whether they received comparable information and resources. Then separate the outcomes: answer quality, total staff time, patient experience and clinical outcomes should not be interchangeable. A promising score can justify a better-designed evaluation without proving a safer care pathway. Look for disagreement handling, missing information and the cost of reviewing the output. These are questions for assessing a proposal, not additional results from the studies above. A useful conclusion should name what was demonstrated and what still needs testing.

Keep these questions handy

  • Did the combined system beat humans alone, AI alone, or both?
  • Were the cases simulated, retrospective, or part of actual prospective care?
  • What information, prompting, interface and training did each group receive?
  • Did the measured benefit include the time and burden of reviewing the output?
  • Was patient benefit measured, or only answer quality and workflow?

Original sources

  1. Goh et al.: Large Language Model Influence on Diagnostic Reasoning (JAMA Network Open, 2024).
  2. Goh et al.: GPT-4 assistance for physician patient-care tasks (Nature Medicine, 2025).
  3. Vaccaro et al.: When combinations of humans and AI are useful (Nature Human Behaviour, 2024).

This AI-assisted reading guide is informational, not medical advice, a systematic review, or a product recommendation. How we work and report corrections.

Follow the evidence as it develops

Read the latest healthcare AI briefing or explore the dated archive. Each edition links to its sources; publication does not guarantee that every claim is correct.

Daily healthcare AI brief

Get the free healthcare AI briefing

One concise briefing built from reporting, research, policy, blogs, videos, and full podcast transcripts.

Free. Sign in to subscribe. Unsubscribe anytime. Already subscribed? Sign in to save and personalize.