Fluent but Wrong: How to Review AI Drafts Before You Sign
AI drafts read like the work of competent colleagues — their most frequent errors are invisible. Five review heuristics against omissions, flipped negations, wrong numbers and misattributions.

Dr. Sven Jungmann
CEO

6:20 p.m., late shift. The attending opens the draft admission note that the new system has generated from the conversation. It is good: cleanly structured, seemingly complete, in better German than some of what humans produce at this hour. She reads it in ninety seconds, changes one word, signs off.
That is what the normal case looks like. It is also the risk case.
The error profile: omissions dominate
Error research on AI-generated clinical text has drawn a consistent pattern over the past two years — with the important caveat that the studies use different denominators and their rates are therefore barely comparable with one another:
- In a comparison of five scribe platforms across 14 standardized encounters (Mayo Clinic Proceedings: Digital Health, 2025), the mean error rate of the generated notes was 26.3 percent; around three quarters of all errors were omissions, and hallucinations were found in 31 percent of the notes [1].
- In a two-month real-world deployment (JMIR Medical Informatics, 2026; 31 physicians, 7,545 generated notes, 356 of them systematically reviewed), omissions were the most frequent error type at 18 percent, ahead of hallucinations (11.5 percent) and inadvertent insertions (9.3 percent) [2]. The denominator is also worth noting: only a fraction of what was generated was reviewed.
- The British CREOLA analysis (npj Digital Medicine, 2025) found, at sentence level, 1.47 percent hallucinations against 3.45 percent omissions — and rated 44 percent of the hallucinations as clinically significant [3].
- In a study of LLM-generated case summaries from the emergency department (UCSF, PLOS Digital Health 2025), 42 percent of the documents contained hallucinations and 47 percent were missing relevant information [4].
The ranking is stable across all designs: omissions dominate, by a factor of two to three. And they are exactly the error class you fundamentally cannot see while reading the text. You cannot notice what is missing as long as you only read what is there. How a generation pipeline can respond to this error profile is described using the physician letter as the example (in German).
The second adversary sits in front of the screen
Human factors research has known the phenomenon since the nineties: automation complacency occurs above all under multitasking load, affects the experienced and beginners alike, and cannot be eliminated by mere practice [5]. Concretely for medicine: in a prescribing study, physicians switched from a correct to an incorrect answer in 5.2 percent of cases after wrong advice from the system [6]. In a radiology vignette experiment, the accuracy of very experienced readers fell from 82 to 45.5 percent under incorrect AI suggestions — experimental conditions, but a clear signal [7].
A fluent text sharpens both. Linguistic quality acts as a signal of competence, and a draft that looks like the work of a good colleague gets read like the work of a good colleague: charitably, briskly, at the level of style. That is precisely why reviewing AI drafts needs a system of its own.
Five review heuristics
- Read against the record, not along the text. Take the structured source first — the diagnosis list, the medication list, open findings, the conversation transcript — and tick off what the draft carries over. Only the comparison finds omissions; reading the draft alone cannot find them, as a matter of principle.
- Handle negations one by one. "No", "not", "without", "discontinued", "paused": allergies, anticoagulation, excluded diagnoses. A flipped negation inverts the meaning while the grammar stays intact. There is a documented case from the older speech-recognition era in which confusing "no vascular flow" with a normal finding ended in an unnecessary procedure [8].
- Every number against its source. Doses, dates, lab values, laterality. Numbers are rare enough in a text that each one can be checked — and none is worth adopting unchecked.
- Check attributions. Who said it: the patient, the relative, a prior report? Speaker misattribution is one of the four documented failure modes of AI scribes, alongside hallucination, omission and contextual misinterpretation [8]. How two correct findings can become an invented diagnosis (in German) shows that even correct individual pieces of information can be combined wrongly.
- Watch your own editing behavior. If over weeks you barely change anything anymore, there are two explanations: the system has become very good — or the review no longer is. The complacency research suggests taking the second explanation seriously, especially with the experienced [5]. At department level, edit rates near zero are a reason for a spot check.
Reviewing can be taught
None of these heuristics costs more than minutes, and all five can be conveyed in an onboarding week. Departments that take this seriously treat reviewing like reading images: as a skill with a system, error classes and practice — with spot checks during chief rounds, a comparison read aloud in the morning briefing, a place in the onboarding curriculum. For those who want to build this in a structured way: our CME-certified course on AI literacy (3 CME points, Ärztekammer Hessen — the Hessian state chamber of physicians) covers exactly this reviewing competence.
In our own conversation documentation we link generated summaries sentence by sentence to the underlying place in the transcript, so that the comparison from heuristic 1 costs no searching; an omission detector that automatically checks drafts against the structured record is planned.
The attending from the late shift needs about four minutes per draft with these five steps. That is the price. It is lower than the other one.
If you want to keep an eye on the evidence on AI documentation: our weekly briefing Visite reviews new studies — concise, referenced, without vendor prose (German; English edition Grand Rounds is in preparation).
Sources
- Anderson TN, Mohan V, Dorr DA, et al. Evaluating the Quality and Safety of Ambient Digital Scribe Platforms Using Simulated Ambulatory Encounters. Mayo Clinic Proceedings: Digital Health. 2025;3(4):100292. doi:10.1016/j.mcpdig.2025.100292
- Taylor SL, et al. Quality of Clinical Notes Created by Ambient Listening Generative AI: Pragmatic Prospective Pilot Study. JMIR Medical Informatics. 2026;14:e86474. doi:10.2196/86474
- Asgari E, et al. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. npj Digital Medicine. 2025;8(1):274. doi:10.1038/s41746-025-01670-7
- Williams CYK, Bains J, Tang T, et al. Evaluating large language models for drafting emergency department encounter summaries. PLOS Digital Health. 2025;4(6):e0000899. doi:10.1371/journal.pdig.0000899
- Parasuraman R, Manzey DH. Complacency and Bias in Human Use of Automation: An Attentional Integration. Human Factors. 2010;52(3):381–410. doi:10.1177/0018720810376055.
- Goddard K, Roudsari A, Wyatt JC. Automation bias: empirical results assessing influencing factors. Int J Med Inform. 2014. https://www.sciencedirect.com/science/article/abs/pii/S1386505614000148
- Dratsch T, et al. Automation Bias in Mammography: The Impact of Artificial Intelligence BI-RADS Suggestions on Reader Performance. Radiology. 2023;307(4):e222176. doi:10.1148/radiol.222176.
- Topaz M, Peltonen LM, Zhang Z. Beyond human ears: navigating the uncharted risks of AI scribes in clinical practice. npj Digital Medicine. 2025. doi:10.1038/s41746-025-01895-6. Ergänzend zur Sprecherverwechslung: npj Digital Medicine. 2025;8:569.
The omission detector mentioned is planned and not yet released; in production are conversation documentation with sentence-level transcript links and basic draft generation.


