AI Physician Letters: What Has to Happen Between Draft and Signature
Peer-reviewed studies show that AI physician letters omit relevant information two to three times more often than they invent it. Which checks belong between draft and signature — and what to measure vendors against.

Dr. Sven Jungmann
CEO

Thursday, 5:40 p.m., an internal medicine ward. Four discharges are planned for tomorrow, four letters still open. The resident opens the drafting tool her hospital has been piloting for a few weeks and, half a minute later, has a draft Arztbrief — the German physician letter — on the screen: salutation correct, clinical course plausibly structured, diagnoses in the right place. The text reads as if written by someone who knows the case.
Precisely this quality is the delicate moment. A fluent, complete-looking text invites skimming. And the most important error class in such drafts is, by its nature, invisible when skimming: it is the sentences that are missing.
Omission beats hallucination — the numbers
There is now solid, peer-reviewed evidence on the error structure of AI-generated clinical text. In 2025, a London group evaluated 450 clinical notes with close to 13,000 sentences in npj Digital Medicine [1]: 1.47 percent of sentences contained hallucinations, that is, invented content; 44 percent of these hallucinations were rated clinically relevant. Omissions stood at 3.45 percent of source sentences — more than twice as frequent.
A study from the University of California, San Francisco on discharge documents from the emergency department shows the same ratio [2]: 42 percent of GPT-4-generated summaries contained hallucinations, 47 percent omitted clinically relevant information.
The public debate about AI physician letters is nonetheless conducted almost exclusively in terms of hallucinations. Whoever only checks whether the letter contains invented content is checking the smaller error class — and the one easier to find. An invented statement is in the text and can give a reader pause. An omitted statement is nowhere. To find it, one would have to know what ought to be in the letter and check every source against the draft. On the fourth letter at 5:40 p.m., nobody does that systematically anymore. How even the merging of two correct findings can go wrong is described elsewhere: How two correct findings become an invented diagnosis (in German).
And yet the drafts really are good. A Freiburg group showed in JMIR Medical Informatics in 2024 that 93.1 percent of German physician letters generated by a locally operated language model were usable with minimal adjustment [3]. And the largest published evaluation of an ambient AI deployment to date — 7,260 physicians at Kaiser Permanente, around 2.5 million patient encounters — found time savings of about 2.1 minutes per encounter, alongside high user satisfaction [4].
Around two minutes per encounter: that is the honest order of magnitude of the time savings. Vendors promising "hours per day" have the published evidence against them. The value of physician-letter AI arises elsewhere — in the completeness of the letter, in the verification costs that today sit invisibly in proofreading and follow-up questions, and in the after-hours documentation that no longer happens.
What has to stand between draft and signature
Legally, the situation is clear: the signed letter is the physician's letter, regardless of who or what wrote the draft. This shifts the question from generation to review — and to whether that review can be evidenced later.
Four steps can be named independently of the vendor.
First, checking against the sources in both directions. "Is what the letter says correct?" is one half. The other: "Does the letter say what the sources say?" This second part — the omission check — is barely feasible manually, because it would require reading every source in full against the draft.
Second, deterministic checks that need no language model: medication reconciled against the documented list, consistency of dates, no codes that exist in no catalogue.
Third, an enforced release: a letter with open questions or missing mandatory entries cannot be signed in the first place — as a property of the software, not as a staff instruction.
Fourth, provability. § 630f of the German Civil Code (BGB) requires that corrections and changes to the documentation remain recognizable [5]. For AI-assisted letters this means: draft, changes and release must be versioned. Who reviewed what and when must still be demonstrable years later — at the latest when the Medizinischer Dienst (MD — the German payers' medical review service) or a court asks. A provable review turns the AI tool into a liability argument in the hospital's favor; an unprovable one turns it into a risk.
Why contextualized generation ages differently from classic text blocks is covered here: Physician letters: text blocks versus contextualized generation (in German).
How aiomics builds the pipeline
At aiomics, base generation is in production today: a competent first draft from the verified, source-linked patient chart in which every field knows its origin. The expansion into the full review pipeline is specified and in development — the following building blocks are all planned and not yet released:
- Omission detector (planned): checks the draft against the structured data of the chart and shows a "not included" panel. Leaving something out remains possible, but becomes a documented decision.
- Dress rehearsal (planned): an adversarial review of the draft from four perspectives — MD auditor, expert reviewer, the practice continuing treatment, the patient — with at most around seven findings per perspective, each with quotation and suggested correction. "No findings" is a permissible result; a review perspective whose findings routinely lead nowhere is dialed down.
- Maker-checker release (planned): finalization is technically blocked as long as open markers remain in the text. Before signature comes a short checklist with substantive confirmations ("discharge dose of apixaban changed — confirm 2×2.5 mg").
- Reviewer diff (planned): for the senior physician, a view of the AI draft against the resident's changes against the current state, with logged signing on behalf ("i.V.").
- Style profile (planned): learns the facility's letter voice from the edit history — stored as readable, editable, deletable German text, never as clinical content.
- Versioning that satisfies § 630f (planned): every version, every review, every release with timestamp and signature facsimile.
What aiomics does not claim: error-free generation. The numbers above apply to language models in general, including ours. The pipeline is the consequence — it is meant to make the physician's review faster and provable; it cannot and must not replace it.
What to measure any vendor against
- Does the system show what it did not include? If a vendor talks only about hallucinations, ask about the omission rate — by the published evidence, it is the error class two to three times larger.
- Does every sentence have a traceable source, and is text without source backing visually distinct, so that attention flows to where the evidence is thin?
- Can a letter with open review items be signed? If yes, the review is a recommendation — and recommendations are the first thing dropped under time pressure.
- How is the physician's review documented? Versioning in the sense of § 630f BGB, timestamps, roles: can it still be demonstrated in five years who reviewed and released the letter?
- What number is the time saving sold with? Minutes per letter are evidence-based. Larger promises should be walked through in detail, number by number.
If you would like to go through these five questions with us — even without evaluating aiomics — write to us. Or simply read along first: our weekly briefing Visite (German; English edition Grand Rounds is in preparation) covers documentation and AI in German healthcare.
Sources
- Asgari E, Montaña-Brown N, Dubois M, et al. A framework to assess clinical safety and hallucination rates of LLMs for medical text summarisation. npj Digital Medicine. 2025;8:274. doi:10.1038/s41746-025-01670-7
- Williams CYK, Bains J, Tang T, et al. Evaluating large language models for drafting emergency department encounter summaries. PLOS Digital Health. 2025;4(6):e0000899. doi:10.1371/journal.pdig.0000899
- Heilmeyer F, Böhringer D, Reinhard T, Arens S, Lyssenko L, Haverkamp C. Viability of open large language models for clinical documentation in German health care: real-world model evaluation study. JMIR Medical Informatics. 2024;12:e59617. doi:10.2196/59617
- Tierney AA, et al. Ambient artificial intelligence scribes: learnings after 1 year and over 2.5 million uses. NEJM Catalyst Innovations in Care Delivery. 2025. doi:10.1056/CAT.25.0040
- § 630f BGB — documentation of treatment. https://www.gesetze-im-internet.de/bgb/__630f.html
The review pipeline described here (omission detector, dress rehearsal, maker-checker release, style profile, reviewer diff, versioning) is planned and not yet released; base draft generation is in production.


