Skip to main content
6 min read

How to Read an AI Scribe Study: Five Questions Before You Believe the Number

Vignette or real-world care? Transcript error or clinically relevant error? Were omissions counted? Five questions that let you place any study on AI documentation within twenty minutes.

Dr. Sven Jungmann

Dr. Sven Jungmann

CEO

Physician reading a study on AI documentation, next to it a vendor brochure with a highlighted percentage

The documents for the procurement round have been circulated. Slide nine of the vendor deck carries a number with an asterisk: "up to 70 percent less documentation time". The asterisk leads to a study in the appendix — twelve pages, English, with confidence intervals. The IT lead asks whether this is solid from a physician's point of view. You have twenty minutes between two appointments.

The good news: twenty minutes are enough. Studies on AI scribes and AI documentation follow recurring patterns, and the places where they hold or break are almost always the same five. They are the questions you would put to any drug study in a journal club — with different vocabulary.

Question 1: vignette or real-world care?

Where did the study take place — on prepared cases or in live operations? The object lesson comes from radiology: 27 radiologists read 50 mammograms to which deliberately wrong AI suggestions had been attached. The accuracy of inexperienced readers fell from around 80 to below 20 percent; that of the very experienced (over 15 years in practice) from 82 to 45.5 percent [1]. A strong finding — and at the same time an experimental vignette design, not real-world care. Vignettes show mechanisms. Transferring effect sizes from vignettes into everyday practice is the most common reading error, in both directions: a scribe that convinces in the quiet demo room has not yet experienced an admission interview with relatives, interposed questions and corridor noise.

For comparison, the gold standard: in the largest randomized study in a real workflow to date — 238 physicians, 14 specialties, two commercial systems against a control group — the better system shortened note-writing time from four and a half minutes to 3:49, a reduction of around ten percent against control [2]. The more robust the design, the more modest the effect. This rule of thumb is the most reliable anchor while reading.

Question 2: what counts as an error — and relative to what?

Modern AI scribes report error rates of about one to three percent; older speech recognition sat at seven to eleven [3]. That sounds like progress, and it is one. Only: what exactly was counted? Some studies define hallucinations narrowly as factual inaccuracies; others count clinical inconsistencies and omissions as well. The rates of different studies are therefore hardly comparable with one another [3]. The denominator changes too: errors per word (the classic word error rate), per note, per encounter. A word error rate of one percent can mean that every fourth note contains a clinically relevant error.

The word error rate also has a blind spot: it measures transcription; whether an error matters clinically, it does not see. A methodological framework in npj Digital Medicine therefore calls for evaluation beyond pure transcription metrics — clinically relevant errors, omissions, workflow integration [4].

Question 3: were omissions counted?

In a two-month pilot with 31 physicians, 356 of 7,545 generated notes were systematically reviewed. Omissions were found in 18 percent of the reviewed notes, making them the most frequent error type, ahead of hallucinations (11.5 percent) and inadvertent insertions (9.3 percent) [5]. Omissions are at the same time the error type reviewers are worst at detecting: you do not see what is missing [3]. A study that reports only hallucinations has measured the smaller error class. How to survey omissions in your own institution, we have described in Testing AI documentation for omissions (in German).

Question 4: who paid — and who chose the endpoints?

The much-cited experience study covering 2.5 million scribe uses at Kaiser Permanente comes from the implementing health system itself [6]. That does not make it worthless — operational data at this scale exists nowhere else, and even this friendliest report in the category arrives, alongside high user satisfaction, at around one minute less after-hours documentation time per appointment [6]. But it is an operator's experience report; an independent proof of effectiveness needs other designs. Read the conflict-of-interest statement of every study you are handed, one by one. And check who chose the endpoints: documentation time and satisfaction are the friendliest endpoints in the category. A commentary in JAMA Network Open points out that central effects — costs, coding intensity, quality of care — have so far not been empirically addressed at all [7].

Question 5: does this hold for your institution?

Most studies in the category come from US outpatient practice, in English, with US documentation logic. A German ward with rounds, handovers and the statutory-insurance form landscape is a different measuring field. German data are scarce: a Freiburg group found in 2024 that 93.1 percent of AI-generated German physician letters were usable with minimal adjustment [9] — what was measured, however, was usability, not error classes. Here too, the question of what was not measured carries. The measurement technique deserves a look as well: one longitudinal study was criticized because the EHR metric used did not capture the time spent in the scribe portal at all — part of the measured saving may only have been displaced time — and because run-in and analysis period were one month each [8]. What transfers, as a rule, are the mechanisms: omissions dominate, automation bias exists, effects are moderate. The effect sizes themselves travel badly.

The five questions on an index card

  1. Setting: vignette, simulation or real workflow?
  2. Error definition: what was counted as an error — and per word, per note or per encounter?
  3. Omissions: surveyed as their own error class, or not at all?
  4. Funding and endpoints: who pays, who measures, which endpoints are missing?
  5. Transferability: language, sector, specialty, measurement window, measurement metric.

If a vendor's study material answers none of the five questions, that is itself an answer.

What remains of "up to 70 percent" after these five questions depends on the individual case. In the most robust available evidence, it is minutes per encounter [2]. Minutes, too, are money and an earlier evening — it is only worth buying them as what they are.

What an evaluation you can trust looks like, we have sorted out in Independent evaluation of clinical AI (in German). If you want to build up the reading technique systematically: our CME-certified course on AI competence (3 CME points, Ärztekammer Hessen — the Hessian state medical chamber) goes deeper into exactly this appraisal method. And in our weekly briefing Visite (German; English edition Grand Rounds is in preparation), we regularly discuss new studies in this category — soberly and with sources.

Sources

  1. Dratsch T, et al. Automation Bias in Mammography: The Impact of Artificial Intelligence BI-RADS Suggestions on Reader Performance. Radiology. 2023;307(4):e222176. doi:10.1148/radiol.222176.
  2. Lukac PJ, Turner W, Vangala S, et al. Ambient AI Scribes in Clinical Practice: A Randomized Trial. NEJM AI. 2025;2(12):AIoa2501000. doi:10.1056/AIoa2501000.
  3. Topaz M, Peltonen LM, Zhang Z. Beyond human ears: navigating the uncharted risks of AI scribes in clinical practice. npj Digital Medicine. 2025;8(1):569. doi:10.1038/s41746-025-01895-6.
  4. Wang H, Yang R, Alwakeel M, et al. An evaluation framework for ambient digital scribing tools in clinical applications. npj Digital Medicine. 2025;8:358. doi:10.1038/s41746-025-01622-1.
  5. Taylor SL, Jost M, MacDonald S, et al. Quality of Clinical Notes Created by Ambient Listening Generative AI: Pragmatic Prospective Pilot Study. JMIR Medical Informatics. 2026;14:e86474. doi:10.2196/86474.
  6. Tierney AA, Gayre G, Hoberman B, et al. Ambient Artificial Intelligence Scribes: Learnings After 1 Year and Over 2.5 Million Uses. NEJM Catalyst Innovations in Care Delivery. 2025;6(5):CAT.25.0040. doi:10.1056/CAT.25.0040.
  7. Ambient AI Scribes — What Is the Return on Investment? JAMA Network Open. 2025. doi:10.1001/jamanetworkopen.2025.53238.
  8. Liu TL, et al. Does AI-Powered Clinical Documentation Enhance Clinician Efficiency? A Longitudinal Study. NEJM AI. 2024;1(12):AIoa2400659. doi:10.1056/AIoa2400659.
  9. Heilmeyer F, Böhringer D, Reinhard T, Arens S, Lyssenko L, Haverkamp C. Viability of Open Large Language Models for Clinical Documentation in German Health Care: Real-World Model Evaluation Study. JMIR Medical Informatics. 2024;12:e59617. doi:10.2196/59617.
#AI scribe study#Ambient scribe evidence#Evaluating AI documentation#Journal club

Keep reading

An executive office at dusk with a packed appointment schedule on screen in the foreground, and a clinician pausing over a chart in a softly lit corridor behind the glass.
Reflections

The Jevons Paradox in Healthcare: Why Faster Doctors Are Not Better Doctors

When AI gives a clinician back ten minutes, the scheduling system tends to fill them with another patient. That instinct quietly converts every efficiency gain into more volume — and mistakes the bottleneck in medicine for time, when it was never time.

Dr. Sven JungmannCEO

This analysis comes from the people behind Visite.

Our weekly newsletter on AI in medicine. Every Friday, rigorously checked.

By signing up you agree to receive Grand Rounds by email. Unsubscribe anytime. More in our privacy policy.

Want to see this in your hospital?

30 minutes. Your questions. Our physician-founder shows you the platform personally.

Book a demo

No commitment. No sales pitch. Physician to physician.