Skip to main content
6 min read

Negation, Laterality, Dose: The Three Most Expensive Detail Errors in AI Documentation

“No vascular flow” becomes a normal finding, left becomes right, 2.5 mg becomes 25: the detail errors with the greatest potential for harm, the evidence on them — and three test routines any hospital can run itself.

Dr. Sven Jungmann

Dr. Sven Jungmann

CEO

Close-up of a draft clinical note in which the words no and left and a dose value are highlighted

The draft admission note is nine-tenths impeccable. History complete, course coherent, language clean. The error sits in a single word: "Metoprolol was discontinued" has become "Metoprolol will be continued". A human does not easily slip that way; a system that turns a half-understood sentence into the most plausible one does.

The clinically most expensive errors of generated documentation are rarely the spectacular ones. They are detail errors in three classes — negations, laterality, numerical values — and what they have in common is that they look plausible and survive fluent proofreading.

In data sheets, these errors rarely appear prominently, because word-based metrics dilute them: at a thousand words per conversation, a flipped "no" disappears into a word error rate of one percent — the metric waters down exactly the errors that are most expensive. That is why it pays to test all three classes deliberately, with a script and a reference. How to do that is included with each class.

Negations: one word carries the entire meaning

Language marks negation sparingly. A "no" in a subordinate clause, a "never", a "discontinued" — if the word gets lost, the statement flips into its opposite, and the sentence remains grammatically impeccable. In German, negation moreover often hides in word formation ("unauffällig" — unremarkable, "beschwerdefrei" — symptom-free, "ausgeschlossen" — ruled out) or arrives late in the sentence. And it travels through chains of reports: a "no evidence of recurrence" that loses its "no" is quoted by every subsequent document that adopts the prior finding. In npj Digital Medicine, a case from the era of older speech recognition is documented in which "no vascular flow" became an unremarkable finding in the transcription; at the end stood an unnecessary procedure [1]. The technology has improved since; the error class has stayed: in an analysis of Whisper transcripts, around one percent contained fabricated passages, 38 percent of those with potential for harm [2].

This is compounded by the speaker question. "Do you have chest pain?" — "No." Whether this no arrives as a patient statement hangs on speaker attribution; in one investigation, the accuracy of role attribution fell from 95 to 82 percent as soon as speaker separation was faulty [3]. How fragile the "who said what" is technically, we have taken apart in AI speech recognition in the hospital: who said what? (in German).

How to test it: write a script with ten negated statements in varying forms — explicit ("no allergies"), elliptical ("Fever? — No."), as a discontinuation ("we stopped taking that in March"), as a negated question with an answer. Record it in a realistic environment and count how many negations arrive unchanged in the draft. Score separately whether a negation is missing or has flipped into its opposite — the second is the more dangerous variant, because the sentence looks complete. Ten out of ten is the standard; anything below wants explaining.

Laterality: left and right are interchangeably plausible

A swapped "left" produces a sentence that sounds exactly as probable as the correct one. In proofreading it only stands out if the reader knows the case — the text itself gives no clue. For AI scribes, according to our research, no robust published rates on laterality errors exist; in the common error taxonomies they run as a subtype of detail errors. Thin data is no all-clear signal here — test it yourself. With document AI, the copy-chain dimension is added: a swapped side in a prior finding is rarely corrected in subsequent letters, because hardly anyone checks against the image again.

How to test it: build side indications systematically into test conversations and dictations — five times left, five times right, plus two switches within the same sentence ("palpable on the left, unremarkable on the right"). For document AI: a finding with multiple, alternating side indications, then a comparison of the extracted fields. Every swap is a finding of the first order.

Numbers: one syllable separates 2.5 from 25

Doses, units, dates. Between 2.5 and 25 milligrams lies acoustically almost nothing and clinically a factor of ten. A comparison across five scribe platforms found an average of 13.9 transcript errors per encounter and a note error rate of 26.3 percent [4]; in the CREOLA work, 44 percent of the hallucinations found were rated clinically significant [5]. Numerical values are among the frequent casualties in both error worlds: they are short and information-dense, and whether 25 is plausible only someone who knows the case can say. The same holds for dates — a shifted admission date can cost more in a billing audit than any stylistic flaw in the letter.

How to test it: dictate a medication list with two dose changes, one of them with a decimal figure and one with a unit switch (milligrams to micrograms), plus two dates in the same sentence. Compare the draft word by word against the script. For document AI, the same applies with a scanned medication list of middling paper quality.

The common denominator — and the test plan

All three classes evade normal proofreading because the erroneous text remains fluent and plausible. The three scripts together can be recorded and scored in an afternoon. Repeat them after every major update of the system — model changes reshuffle error profiles, and the test from spring proves nothing about the version from autumn. How the three test routines fit into a structured trial — with sample, counting rules and stop criteria — is in Trialing clinical AI: the 14-day test protocol (in German).

In the aiomics conversation documentation, this is the point where we presuppose the least trust: negating segments on medication, allergies and diagnoses are displayed inline with the original quote from the transcript, the preservation of negations is measured as a dedicated benchmark, and speaker attributions below 85 percent confidence are marked as uncertain. This is in production — deliberately cut narrowly to the admission interview in a one-on-one setting.

If you want to run the three test routines once in your own setting, write to us — even without evaluating aiomics. In our weekly briefing Visite (German; English edition Grand Rounds is in preparation), we regularly collect what the evidence on AI documentation yields, detail errors included.

Sources

  1. Topaz M, Peltonen LM, Zhang Z. Beyond human ears: navigating the uncharted risks of AI scribes in clinical practice. npj Digital Medicine. 2025;8(1):569. doi:10.1038/s41746-025-01895-6.
  2. Koenecke A, Choi ASG, Mei KX, Schellmann H, Sloane M. Careless Whisper: Speech-to-Text Hallucination Harms. In: Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT '24). New York: ACM; 2024. p. 1672–1681. doi:10.1145/3630106.3658996.
  3. Untersuchung zur Rollenzuordnung in klinischen Gesprächstranskripten (BERT-basiert). https://pmc.ncbi.nlm.nih.gov/articles/PMC12632672/.
  4. Anderson TN, Mohan V, Dorr DA, Ratwani RM, Biro JM, Gold JA. Evaluating the Quality and Safety of Ambient Digital Scribe Platforms Using Simulated Ambulatory Encounters. Mayo Clinic Proceedings: Digital Health. 2025;3(4):100292. doi:10.1016/j.mcpdig.2025.100292.
  5. Asgari E, Montaña-Brown N, Dubois M, et al. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. npj Digital Medicine. 2025;8(1):274. doi:10.1038/s41746-025-01670-7.
#AI documentation errors#Negation speech recognition#Medication errors AI#Ambient scribe risks

The aiomics conversation documentation described here is in production for admission interviews in one-on-one settings; extensions to further conversation types are planned.

Keep reading

An executive office at dusk with a packed appointment schedule on screen in the foreground, and a clinician pausing over a chart in a softly lit corridor behind the glass.
Reflections

The Jevons Paradox in Healthcare: Why Faster Doctors Are Not Better Doctors

When AI gives a clinician back ten minutes, the scheduling system tends to fill them with another patient. That instinct quietly converts every efficiency gain into more volume — and mistakes the bottleneck in medicine for time, when it was never time.

Dr. Sven JungmannCEO
Draft physician letter on a screen next to a stack of source documents; a highlighted finding in the stack is missing from the draft

Omissions: The Error Class You Cannot See When Proofreading

AI drafts leave something out two to three times more often than they invent something — and omissions leave no trace in the text. How to measure the dominant error class in your own institution, in five steps.

Dr. Sven JungmannCEO

This analysis comes from the people behind Visite.

Our weekly newsletter on AI in medicine. Every Friday, rigorously checked.

By signing up you agree to receive Grand Rounds by email. Unsubscribe anytime. More in our privacy policy.

Want to see this in your hospital?

30 minutes. Your questions. Our physician-founder shows you the platform personally.

Book a demo

No commitment. No sales pitch. Physician to physician.