Skip to main content
6 min read

Automation Bias at the Bedside: Why Edit Rates Near Zero Are a Warning Sign

Automation complacency affects the experienced and the novice alike, grows under load, and does not disappear with practice. What human-factors research means for everyday ward work with AI drafts — and why a correction rate near zero is not proof of quality.

Dr. Sven Jungmann

Dr. Sven Jungmann

CEO

Physician at a ward computer late in the evening, a stack of already approved AI drafts, the mouse pointer hovering over the approve button

9:40 p.m., ward office. The fourteenth AI draft of the day is waiting for sign-off. The first thirteen were good — two small things, otherwise clean. The resident reads the fourteenth faster than the first. His brain has learned, over thirteen rounds, that these texts are correct. Tomorrow there will be sixteen drafts, and the duty roster is the same.

Human-factors research has known this mechanism for decades — from cockpits, control rooms and radar stations, long before there were language models. It has also measured how it behaves. For departments introducing AI drafts into their documentation, a look into this literature is worth more than any product data sheet.

What the research actually shows

The terms come from a classic 1997 paper: "misuse" is over-reliance on automation — the monitoring falls asleep; "disuse" is the counterpart, the neglect of a system, usually after too many false alarms [1]. A 2010 review, co-authored by Dietrich Manzey of TU Berlin, integrates the experimental evidence into three core findings [2].

First: automation complacency arises above all under multiple workloads, when manual tasks compete for attention with the automation being monitored. That is not an exotic laboratory condition. It is the description of a ward shift.

Second: it affects novices and experts alike. Experience does not protect.

Third: it cannot be eliminated by mere practice. "People will get used to checking" is not a strategy but a hope.

The flip side is also part of the literature: systems that warn too often without cause produce disuse — they get ignored or switched off [1]. A review regime that treats every AI line with maximum distrust is therefore just as unstable as blind trust. Attention is a finite, manageable resource [2]; the question is where it is directed.

From the laboratory to the clinic

For healthcare, the effect is quantified. A systematic review of decision support systems reached a double finding: such systems improve overall performance — and at the same time introduce new, often unrecorded errors [3]. A follow-up study by the same group got more concrete: in 5.2 percent of the prescribing cases examined, physicians switched from a correct to an incorrect decision after a wrong system suggestion [4]. The influencing factors identified were trust in the system, one's own decision confidence and the difficulty of the task [4].

Add to this the 2023 mammography study, by now the object lesson of the field: 27 radiologists, 50 readings, deliberately wrong AI suggestions. The inexperienced fell from around 80 to below 20 percent accuracy, the very experienced with over 15 years in practice from 82 to 45.5 percent [5]. That was a vignette experiment, not a real-world care study — but it answers the question of whether seniority immunizes. It roughly halves the damage. It does not prevent it.

The paradox of good systems

The more reliable a system, the more rarely the check is rewarded — and the faster it withers. A tool with many errors keeps its reviewers awake; a tool with few, rare errors trains them out of the habit. The rare errors then meet the weakest review. Modern documentation AI sits exactly in this zone: good enough to generate trust, and far from error-free — in a comparison across five platforms, the note error rate was 26.3 percent [6]. For the fourteenth draft of the evening, this means: the probability of an error is real — what has fallen is only the probability of finding it.

An observation from the field fits this: in a two-month pilot, 31 physicians generated 7,545 ambient notes; 356 of them were systematically reviewed [7]. The study does not explain why the rest went unreviewed, and one should not read more into it than is there. But the ratio probably describes normal operations more honestly than any self-report on review discipline.

The metric leadership should see

From this follows a look at an inconspicuous number: the edit rate. How often do your physicians actually change something in AI drafts? A correction rate near zero has two possible explanations. Either the system is practically error-free — a state the literature does not know for this tool class [6]. Or the review has gone out. The second explanation is the more likely one, and unlike the first, it can be managed:

  • Collect edit rates in aggregate and discuss them in the department — as a process metric at team level, without reference to individuals.
  • Institutionalize spot-check second review: each month, a fixed number of notes checked against the sources, rotating within the team.
  • Occasionally slip in a prepared case with a known error. If it is found, the review is alive. If it is waved through, you know in time — and without patient harm.
  • Budget review time as working time. The strongest driver of complacency is competing load [2]; whoever does not budget review time has de facto abolished it.

A validated target value for edit rates does not exist, and that honesty is part of it. The metric's value lies in its movement: a department whose correction rate falls over months while nothing about the system has changed is watching its own trust grow — and should re-measure whether that trust is earned. The most frequent error type, the omission, remains structurally invisible to mere proofreading anyway [7]; the spot check therefore tests against the sources and never relies on the reading impression alone.

What to look for in the proofreading itself — omissions, negations, numbers, attributions — we have compiled in Fluent, but wrong: reviewing AI texts (in German). A structured 14-day protocol for the trial phase is in Trialing clinical AI (in German).

We have turned this literature into a design principle for our own modules: at aiomics, an edit rate near zero counts as a review signal that triggers a workflow review — and for the planned batch processing it is laid down that sustained blind acceptance above 99 percent automatically triggers a human-factors review.

If you want to anchor the topic in your department: our CME-certified course on AI competence (3 CME points, Ärztekammer Hessen — the Hessian state medical chamber) goes deeper into exactly this reviewing competence. And in our weekly briefing Visite (German; English edition Grand Rounds is in preparation), we write regularly about the evidence behind AI documentation — including the uncomfortable kind.

Sources

  1. Parasuraman R, Riley V. Humans and Automation: Use, Misuse, Disuse, Abuse. Human Factors. 1997;39(2):230–253. doi:10.1518/001872097778543886.
  2. Parasuraman R, Manzey DH. Complacency and Bias in Human Use of Automation: An Attentional Integration. Human Factors. 2010;52(3):381–410. doi:10.1177/0018720810376055.
  3. Goddard K, Roudsari A, Wyatt JC. Automation bias: a systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association. 2012;19(1):121–127. doi:10.1136/amiajnl-2011-000089.
  4. Goddard K, Roudsari A, Wyatt JC. Automation bias: empirical results assessing influencing factors. International Journal of Medical Informatics. 2014;83(5):368–375. doi:10.1016/j.ijmedinf.2014.01.001.
  5. Dratsch T, Chen X, Rezazade Mehrizi M, et al. Automation Bias in Mammography: The Impact of Artificial Intelligence BI-RADS Suggestions on Reader Performance. Radiology. 2023;307(4):e222176. doi:10.1148/radiol.222176.
  6. Anderson TN, Mohan V, Dorr DA, Ratwani RM, Biro JM, Gold JA. Evaluating the Quality and Safety of Ambient Digital Scribe Platforms Using Simulated Ambulatory Encounters. Mayo Clinic Proceedings: Digital Health. 2025;3(4):100292. doi:10.1016/j.mcpdig.2025.100292.
  7. Taylor SL, Jost M, MacDonald S, et al. Quality of Clinical Notes Created by Ambient Listening Generative AI: Pragmatic Prospective Pilot Study. JMIR Medical Informatics. 2026;14:e86474. doi:10.2196/86474.
#Automation bias#AI hospital safety#Human factors medicine#Reviewing AI drafts

The human-factors trigger for batch processing described here belongs to a planned module that is not yet in production.

Keep reading

Editorial collage: a quarterly report whose figures are connected by threads to queries and control charts, a watermark across the page

Why aiomics for QM reports and quality analytics

A QM report in which every figure resolves to its query and the language model never sees individual data. All of it is planned — and we are laying the architecture open now anyway, so you can measure us against it.

Dr. Sven JungmannCEO
Editorial collage: a §301 data record as a form grid, every field connected to its source passage by threads; a traffic light shows the formal check status

Why aiomics for coding suggestions and §301 preparation

Coding suggestions only for what is explicitly documented, a deterministic §301 gate before transmission: what we are building, why it is designed this way — and why we disclose the status before the module is released.

Dr. Sven JungmannCEO
Editorial collage: a physician letter whose sentences are connected to source documents by threads; a marked gap shows a missing entry

Why aiomics for discharge letters and physician letters

Draft generation is live, the verified pipeline is planned: how aiomics builds physician and discharge letters with a source pointer per sentence, why omissions are the bigger error — and where we are not the right choice.

Dr. Sven JungmannCEO

This analysis comes from the people behind Visite.

Our weekly newsletter on AI in medicine. Every Friday, rigorously checked.

By signing up you agree to receive Grand Rounds by email. Unsubscribe anytime. More in our privacy policy.

Want to see this in your hospital?

30 minutes. Your questions. Our physician-founder shows you the platform personally.

Book a demo

No commitment. No sales pitch. Physician to physician.