Ambient AI scribes: how to read the evidence

A faster draft is not the same as a safer note, and a happier clinician is not the same as a measured time saving. When an ambient AI announcement arrives, start by identifying which of those claims was actually tested.

What was compared?

A randomized trial of 238 outpatient physicians compared DAX Copilot, Nabla, and usual care. Nabla reduced time-in-note relative to control; DAX did not show a statistically significant reduction on that outcome. The authors describe potential improvements in secondary well-being measures as needing confirmation. This is evidence about specific products, participants, and a study period, not a verdict on every ambient scribe or its current version. [1]

Which kind of improvement?

Another randomized implementation study found reductions in work exhaustion and interpersonal disengagement without a significant increase in professional fulfillment. Those are different outcomes. Our reading rule: keep the measure attached to the claim. Do not turn a change on one questionnaire into an assertion that burnout was eliminated, staffing costs fell, or patient outcomes improved. [2]

Where did the work go?

Look for adoption and editing burden, not just whether a tool was available. In the 238-physician trial, clinicians reported occasional clinically significant inaccuracies. A polished note therefore does not establish factual reliability. For a local evaluation, ask how corrections are captured, who reviews the final note, and whether the workflow measures review time as well as drafting time. These are OpenRounds' evaluation questions, not findings that every product has met. [1]

What would change your decision?

Before comparing vendors, write down the task, the people who would use the tool, and the outcome that would make a trial worthwhile. Keep an efficiency question separate from a documentation-quality question. A podcast can explain how a team introduced a tool or trained staff; it cannot, by itself, establish that a measured benefit will transfer to your setting. Follow the original study rather than relying on an announcement's headline.

Keep these questions handy

  • Was this a randomized comparison, a before-and-after rollout, or a testimonial?
  • Does the reported result concern time, experience, accuracy, or patient outcomes?
  • Which users and encounters actually used the tool, and who corrected its output?
  • Does the study identify the product version, setting, comparator, and follow-up period?

Original sources

  1. Lukac et al. Ambient AI Scribes in Clinical Practice: A Randomized Trial. NEJM AI (2025).
  2. Afshar et al. A Pragmatic Randomized Controlled Trial of Ambient Artificial Intelligence to Improve Health Practitioner Well-Being. NEJM AI (2025).

This AI-assisted reading guide is informational, not medical advice, a systematic review, or a product recommendation. How we work and report corrections.

Follow the evidence as it develops

Read the latest healthcare AI briefing or explore the dated archive. Each edition links to its sources; publication does not guarantee that every claim is correct.

Daily healthcare AI brief

Get the free healthcare AI briefing

One concise briefing built from reporting, research, policy, blogs, videos, and full podcast transcripts.

Free. No account needed. Unsubscribe anytime. Already subscribed? Sign in to save and personalize.