Automation Bias at the Bedside: Why Edit Rates Near Zero Are a Warning Sign
Automation complacency affects the experienced and the novice alike, grows under load, and does not disappear with practice. What human-factors research means for everyday ward work with AI drafts — and why a correction rate near zero is not proof of quality.

Dr. Sven Jungmann
CEO

9:40 p.m., ward office. The fourteenth AI draft of the day is waiting for sign-off. The first thirteen were good — two small things, otherwise clean. The resident reads the fourteenth faster than the first. His brain has learned, over thirteen rounds, that these texts are correct. Tomorrow there will be sixteen drafts, and the duty roster is the same.
Human-factors research has known this mechanism for decades — from cockpits, control rooms and radar stations, long before there were language models. It has also measured how it behaves. For departments introducing AI drafts into their documentation, a look into this literature is worth more than any product data sheet.
What the research actually shows
The terms come from a classic 1997 paper: "misuse" is over-reliance on automation — the monitoring falls asleep; "disuse" is the counterpart, the neglect of a system, usually after too many false alarms [1]. A 2010 review, co-authored by Dietrich Manzey of TU Berlin, integrates the experimental evidence into three core findings [2].
First: automation complacency arises above all under multiple workloads, when manual tasks compete for attention with the automation being monitored. That is not an exotic laboratory condition. It is the description of a ward shift.
Second: it affects novices and experts alike. Experience does not protect.
Third: it cannot be eliminated by mere practice. "People will get used to checking" is not a strategy but a hope.
The flip side is also part of the literature: systems that warn too often without cause produce disuse — they get ignored or switched off [1]. A review regime that treats every AI line with maximum distrust is therefore just as unstable as blind trust. Attention is a finite, manageable resource [2]; the question is where it is directed.
From the laboratory to the clinic
For healthcare, the effect is quantified. A systematic review of decision support systems reached a double finding: such systems improve overall performance — and at the same time introduce new, often unrecorded errors [3]. A follow-up study by the same group got more concrete: in 5.2 percent of the prescribing cases examined, physicians switched from a correct to an incorrect decision after a wrong system suggestion [4]. The influencing factors identified were trust in the system, one's own decision confidence and the difficulty of the task [4].
Add to this the 2023 mammography study, by now the object lesson of the field: 27 radiologists, 50 readings, deliberately wrong AI suggestions. The inexperienced fell from around 80 to below 20 percent accuracy, the very experienced with over 15 years in practice from 82 to 45.5 percent [5]. That was a vignette experiment, not a real-world care study — but it answers the question of whether seniority immunizes. It roughly halves the damage. It does not prevent it.
The paradox of good systems
The more reliable a system, the more rarely the check is rewarded — and the faster it withers. A tool with many errors keeps its reviewers awake; a tool with few, rare errors trains them out of the habit. The rare errors then meet the weakest review. Modern documentation AI sits exactly in this zone: good enough to generate trust, and far from error-free — in a comparison across five platforms, the note error rate was 26.3 percent [6]. For the fourteenth draft of the evening, this means: the probability of an error is real — what has fallen is only the probability of finding it.
An observation from the field fits this: in a two-month pilot, 31 physicians generated 7,545 ambient notes; 356 of them were systematically reviewed [7]. The study does not explain why the rest went unreviewed, and one should not read more into it than is there. But the ratio probably describes normal operations more honestly than any self-report on review discipline.
The metric leadership should see
From this follows a look at an inconspicuous number: the edit rate. How often do your physicians actually change something in AI drafts? A correction rate near zero has two possible explanations. Either the system is practically error-free — a state the literature does not know for this tool class [6]. Or the review has gone out. The second explanation is the more likely one, and unlike the first, it can be managed:
- Collect edit rates in aggregate and discuss them in the department — as a process metric at team level, without reference to individuals.
- Institutionalize spot-check second review: each month, a fixed number of notes checked against the sources, rotating within the team.
- Occasionally slip in a prepared case with a known error. If it is found, the review is alive. If it is waved through, you know in time — and without patient harm.
- Budget review time as working time. The strongest driver of complacency is competing load [2]; whoever does not budget review time has de facto abolished it.
A validated target value for edit rates does not exist, and that honesty is part of it. The metric's value lies in its movement: a department whose correction rate falls over months while nothing about the system has changed is watching its own trust grow — and should re-measure whether that trust is earned. The most frequent error type, the omission, remains structurally invisible to mere proofreading anyway [7]; the spot check therefore tests against the sources and never relies on the reading impression alone.
What to look for in the proofreading itself — omissions, negations, numbers, attributions — we have compiled in Fluent, but wrong: reviewing AI texts (in German). A structured 14-day protocol for the trial phase is in Trialing clinical AI (in German).
We have turned this literature into a design principle for our own modules: at aiomics, an edit rate near zero counts as a review signal that triggers a workflow review — and for the planned batch processing it is laid down that sustained blind acceptance above 99 percent automatically triggers a human-factors review.
If you want to anchor the topic in your department: our CME-certified course on AI competence (3 CME points, Ärztekammer Hessen — the Hessian state medical chamber) goes deeper into exactly this reviewing competence. And in our weekly briefing Visite (German; English edition Grand Rounds is in preparation), we write regularly about the evidence behind AI documentation — including the uncomfortable kind.
Sources
- Parasuraman R, Riley V. Humans and Automation: Use, Misuse, Disuse, Abuse. Human Factors. 1997;39(2):230–253. doi:10.1518/001872097778543886.
- Parasuraman R, Manzey DH. Complacency and Bias in Human Use of Automation: An Attentional Integration. Human Factors. 2010;52(3):381–410. doi:10.1177/0018720810376055.
- Goddard K, Roudsari A, Wyatt JC. Automation bias: a systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association. 2012;19(1):121–127. doi:10.1136/amiajnl-2011-000089.
- Goddard K, Roudsari A, Wyatt JC. Automation bias: empirical results assessing influencing factors. International Journal of Medical Informatics. 2014;83(5):368–375. doi:10.1016/j.ijmedinf.2014.01.001.
- Dratsch T, Chen X, Rezazade Mehrizi M, et al. Automation Bias in Mammography: The Impact of Artificial Intelligence BI-RADS Suggestions on Reader Performance. Radiology. 2023;307(4):e222176. doi:10.1148/radiol.222176.
- Anderson TN, Mohan V, Dorr DA, Ratwani RM, Biro JM, Gold JA. Evaluating the Quality and Safety of Ambient Digital Scribe Platforms Using Simulated Ambulatory Encounters. Mayo Clinic Proceedings: Digital Health. 2025;3(4):100292. doi:10.1016/j.mcpdig.2025.100292.
- Taylor SL, Jost M, MacDonald S, et al. Quality of Clinical Notes Created by Ambient Listening Generative AI: Pragmatic Prospective Pilot Study. JMIR Medical Informatics. 2026;14:e86474. doi:10.2196/86474.
The human-factors trigger for batch processing described here belongs to a planned module that is not yet in production.


