Negation, Laterality, Dose: The Three Most Expensive Detail Errors in AI Documentation
“No vascular flow” becomes a normal finding, left becomes right, 2.5 mg becomes 25: the detail errors with the greatest potential for harm, the evidence on them — and three test routines any hospital can run itself.

Dr. Sven Jungmann
CEO

The draft admission note is nine-tenths impeccable. History complete, course coherent, language clean. The error sits in a single word: "Metoprolol was discontinued" has become "Metoprolol will be continued". A human does not easily slip that way; a system that turns a half-understood sentence into the most plausible one does.
The clinically most expensive errors of generated documentation are rarely the spectacular ones. They are detail errors in three classes — negations, laterality, numerical values — and what they have in common is that they look plausible and survive fluent proofreading.
In data sheets, these errors rarely appear prominently, because word-based metrics dilute them: at a thousand words per conversation, a flipped "no" disappears into a word error rate of one percent — the metric waters down exactly the errors that are most expensive. That is why it pays to test all three classes deliberately, with a script and a reference. How to do that is included with each class.
Negations: one word carries the entire meaning
Language marks negation sparingly. A "no" in a subordinate clause, a "never", a "discontinued" — if the word gets lost, the statement flips into its opposite, and the sentence remains grammatically impeccable. In German, negation moreover often hides in word formation ("unauffällig" — unremarkable, "beschwerdefrei" — symptom-free, "ausgeschlossen" — ruled out) or arrives late in the sentence. And it travels through chains of reports: a "no evidence of recurrence" that loses its "no" is quoted by every subsequent document that adopts the prior finding. In npj Digital Medicine, a case from the era of older speech recognition is documented in which "no vascular flow" became an unremarkable finding in the transcription; at the end stood an unnecessary procedure [1]. The technology has improved since; the error class has stayed: in an analysis of Whisper transcripts, around one percent contained fabricated passages, 38 percent of those with potential for harm [2].
This is compounded by the speaker question. "Do you have chest pain?" — "No." Whether this no arrives as a patient statement hangs on speaker attribution; in one investigation, the accuracy of role attribution fell from 95 to 82 percent as soon as speaker separation was faulty [3]. How fragile the "who said what" is technically, we have taken apart in AI speech recognition in the hospital: who said what? (in German).
How to test it: write a script with ten negated statements in varying forms — explicit ("no allergies"), elliptical ("Fever? — No."), as a discontinuation ("we stopped taking that in March"), as a negated question with an answer. Record it in a realistic environment and count how many negations arrive unchanged in the draft. Score separately whether a negation is missing or has flipped into its opposite — the second is the more dangerous variant, because the sentence looks complete. Ten out of ten is the standard; anything below wants explaining.
Laterality: left and right are interchangeably plausible
A swapped "left" produces a sentence that sounds exactly as probable as the correct one. In proofreading it only stands out if the reader knows the case — the text itself gives no clue. For AI scribes, according to our research, no robust published rates on laterality errors exist; in the common error taxonomies they run as a subtype of detail errors. Thin data is no all-clear signal here — test it yourself. With document AI, the copy-chain dimension is added: a swapped side in a prior finding is rarely corrected in subsequent letters, because hardly anyone checks against the image again.
How to test it: build side indications systematically into test conversations and dictations — five times left, five times right, plus two switches within the same sentence ("palpable on the left, unremarkable on the right"). For document AI: a finding with multiple, alternating side indications, then a comparison of the extracted fields. Every swap is a finding of the first order.
Numbers: one syllable separates 2.5 from 25
Doses, units, dates. Between 2.5 and 25 milligrams lies acoustically almost nothing and clinically a factor of ten. A comparison across five scribe platforms found an average of 13.9 transcript errors per encounter and a note error rate of 26.3 percent [4]; in the CREOLA work, 44 percent of the hallucinations found were rated clinically significant [5]. Numerical values are among the frequent casualties in both error worlds: they are short and information-dense, and whether 25 is plausible only someone who knows the case can say. The same holds for dates — a shifted admission date can cost more in a billing audit than any stylistic flaw in the letter.
How to test it: dictate a medication list with two dose changes, one of them with a decimal figure and one with a unit switch (milligrams to micrograms), plus two dates in the same sentence. Compare the draft word by word against the script. For document AI, the same applies with a scanned medication list of middling paper quality.
The common denominator — and the test plan
All three classes evade normal proofreading because the erroneous text remains fluent and plausible. The three scripts together can be recorded and scored in an afternoon. Repeat them after every major update of the system — model changes reshuffle error profiles, and the test from spring proves nothing about the version from autumn. How the three test routines fit into a structured trial — with sample, counting rules and stop criteria — is in Trialing clinical AI: the 14-day test protocol (in German).
In the aiomics conversation documentation, this is the point where we presuppose the least trust: negating segments on medication, allergies and diagnoses are displayed inline with the original quote from the transcript, the preservation of negations is measured as a dedicated benchmark, and speaker attributions below 85 percent confidence are marked as uncertain. This is in production — deliberately cut narrowly to the admission interview in a one-on-one setting.
If you want to run the three test routines once in your own setting, write to us — even without evaluating aiomics. In our weekly briefing Visite (German; English edition Grand Rounds is in preparation), we regularly collect what the evidence on AI documentation yields, detail errors included.
Sources
- Topaz M, Peltonen LM, Zhang Z. Beyond human ears: navigating the uncharted risks of AI scribes in clinical practice. npj Digital Medicine. 2025;8(1):569. doi:10.1038/s41746-025-01895-6.
- Koenecke A, Choi ASG, Mei KX, Schellmann H, Sloane M. Careless Whisper: Speech-to-Text Hallucination Harms. In: Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT '24). New York: ACM; 2024. p. 1672–1681. doi:10.1145/3630106.3658996.
- Untersuchung zur Rollenzuordnung in klinischen Gesprächstranskripten (BERT-basiert). https://pmc.ncbi.nlm.nih.gov/articles/PMC12632672/.
- Anderson TN, Mohan V, Dorr DA, Ratwani RM, Biro JM, Gold JA. Evaluating the Quality and Safety of Ambient Digital Scribe Platforms Using Simulated Ambulatory Encounters. Mayo Clinic Proceedings: Digital Health. 2025;3(4):100292. doi:10.1016/j.mcpdig.2025.100292.
- Asgari E, Montaña-Brown N, Dubois M, et al. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. npj Digital Medicine. 2025;8(1):274. doi:10.1038/s41746-025-01670-7.
The aiomics conversation documentation described here is in production for admission interviews in one-on-one settings; extensions to further conversation types are planned.


