Omissions: The Error Class You Cannot See When Proofreading
AI drafts leave something out two to three times more often than they invent something — and omissions leave no trace in the text. How to measure the dominant error class in your own institution, in five steps.

Dr. Sven Jungmann
CEO

The draft reads well. The senior physician goes through it paragraph by paragraph: reason for admission correct, course correct, medication correct, the plan is plausible. She changes two formulations and approves. What she could not see: the admission labs contained a borderline creatinine value, the prior findings a drug intolerance. Both are in the source documents. Neither is anywhere in the letter — and nothing in the letter suggests that something is missing.
That is the asymmetry between the two large error classes of generated documentation. An invented statement is in the text; an attentive reader can stumble over it. An omission leaves no trace in the document being reviewed. How two correct findings can still produce a wrong record, we have described in How two correct findings become an invented diagnosis (in German). Here we deal with the reverse, more frequent case: something is missing.
The evidence: omissions dominate
The findings on this are strikingly consistent across research groups and methodologies.
- In the CREOLA work on LLM-generated clinical documentation, the hallucination rate was 1.47 percent, the omission rate 3.45 percent — more than double. 44 percent of the hallucinations found were rated clinically significant [1].
- At the University of California San Francisco, 42 percent of GPT-4-generated discharge documents from the emergency department contained hallucinations — and 47 percent missing relevant information [2].
- A comparison of five scribe platforms across 14 encounters each found a note error rate of 26.3 percent; around three quarters of all errors were omissions [3].
- In a two-month pilot (31 physicians, 7,545 generated notes, 356 of them systematically reviewed), omissions were the most frequent error type at 18 percent of the reviewed notes, ahead of hallucinations at 11.5 and inadvertent insertions at 9.3 percent [4].
One caveat belongs here: the studies measure differently — sometimes per sentence, sometimes per note, sometimes as a share of all errors found — and their rates can hardly be compared with one another [5]. The pattern is stable nonetheless: something is missing more often than something is invented. Whoever aligns their review regime solely to hallucinations is reviewing the smaller error class.
Why proofreading alone is not enough
Proofreading checks what is on the page. For omissions, that is precisely the blind spot: in the literature, they are considered the error type hardest for reviewers to detect, because the text gives no occasion to pause [5]. A draft can appear complete and still withhold the one finding on which follow-up treatment or billing will later hang.
The review therefore needs a second list: what ought to be on the page. This list can be generated — from the source documents, before reading the draft.
Editing or omission?
One distinction up front, because the entire measurement hangs on it: leaving something out is not per se an error. A good letter selects. General practitioners explicitly want prioritized, shorter letters — with actionable to-do lists, flagged incidental findings and substantiated medication changes, expressly without exhaustive data reproduction [6]. The difference lies in the decision: editing is a deliberate judgment about relevance. An omission in the sense measured here is a leaving-out that nobody decided — and that nobody, therefore, can answer for.
Measuring omissions in your own institution: five steps
You need no study design and no ethics committee for this: it is an internal quality review of documents that pass through your sign-off anyway.
- Draw a sample. 20 to 30 cases over two weeks, spread across different authors, times of day and case types. Routine cases included; the sample should reflect everyday practice, not a collection of rarities. If several document types are in use — physician letter, admission note, conversation summary — review them separately: error profiles differ by document type.
- Create the source list. Before looking at the draft, one reviewer extracts the review-relevant items from the source documents: diagnoses, medication with dose, allergies and intolerances, pathological and borderline findings, pending results, follow-up obligations.
- Compare. A second person checks the draft against the source list and marks every item: adopted, left out with reason, or missing. Whoever created the source list does not review the same case. Contested items go into a short consensus discussion; also note how often the reviewers disagreed — that rate says something about the sharpness of your counting rule.
- Count. An omission counts if an item is relevant for follow-up treatment or billing and neither appears in the draft nor is recognizably absent for a stated reason. What is reported is the share of cases with at least one relevant omission — per case, so the figure remains comparable with the literature.
- Interpret. Values in the range of the published studies — depending on definition, between just under a fifth and half of the documents [2,4] — are not an outlier but the normal state of this tool class. The real question is what your review process catches of it.
The measurement works as a baseline before any procurement decision — and as a repeat measurement in live operations once a system has long been introduced.
What to measure vendors against
- Does the vendor report an omission rate — or exclusively hallucination figures?
- Against which reference was it measured: transcript, source documents, or a gold-standard note? Omissions are only measurable against the sources.
- Does the system show what it left out, or does the reviewing eye have to find it by itself?
- Can the five-step measurement be run on your own material during a trial? A structured protocol for that is in Trialing clinical AI: the 14-day test protocol (in German).
In our own document pipeline, we are planning an omission detector for this problem: it checks every draft against the structured data of the verified record and makes what was left out visible in a "Not adopted" panel, so that leaving out becomes a logged decision — the feature is in development and not yet in production.
If you want to run the five-step measurement once on your own documents, write to us — even without evaluating aiomics. In our weekly briefing Visite (German; English edition Grand Rounds is in preparation), we write regularly about documentation quality and the evidence behind it.
Sources
- Asgari E, Montaña-Brown N, Dubois M, et al. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. npj Digital Medicine. 2025;8(1):274. doi:10.1038/s41746-025-01670-7.
- Williams CYK, Bains J, Tang T, et al. Evaluating large language models for drafting emergency department encounter summaries. PLOS Digital Health. 2025;4(6):e0000899. doi:10.1371/journal.pdig.0000899.
- Anderson TN, Mohan V, Dorr DA, Ratwani RM, Biro JM, Gold JA. Evaluating the Quality and Safety of Ambient Digital Scribe Platforms Using Simulated Ambulatory Encounters. Mayo Clinic Proceedings: Digital Health. 2025;3(4):100292. doi:10.1016/j.mcpdig.2025.100292.
- Taylor SL, Jost M, MacDonald S, et al. Quality of Clinical Notes Created by Ambient Listening Generative AI: Pragmatic Prospective Pilot Study. JMIR Medical Informatics. 2026;14:e86474. doi:10.2196/86474.
- Topaz M, Peltonen LM, Zhang Z. Beyond human ears: navigating the uncharted risks of AI scribes in clinical practice. npj Digital Medicine. 2025;8(1):569. doi:10.1038/s41746-025-01895-6.
- Primary Care Physicians' Perspectives on High-Quality Discharge Summaries. Journal of General Internal Medicine. 2023. doi:10.1007/s11606-023-08541-5.
The aiomics omission detector mentioned in the text is in development and not yet in production.


