Skip to main content
6 min read

Trialing Clinical AI: The 14-Day Test Protocol

Four speakers in the room, flipped negations, corridor noise, third-generation faxes: a vendor-neutral 14-day protocol for trialing clinical AI — with scenarios, counting rules and stop criteria.

Dr. Sven Jungmann

Dr. Sven Jungmann

CEO

Clipboard with a test protocol and tally sheet next to a hospital computer, an admission interview with several people in the background

The test deployment starts on Monday. In the demo, the system was convincing: quiet room, one voice, a cleanly scanned physician letter. Your hospital sounds different. Admission interviews with relatives, questions cutting across sentences, monitors, corridor traffic — and faxes that already have two copying generations behind them.

Whether a system can handle that is not decided in the demo but in a structured trial. The following protocol is vendor-neutral and can be run in 14 days. It needs two physicians, one coordinating staff member and a spreadsheet.

Week 0: preparation

  1. Name the review team. Two physician reviewers of different experience levels, plus one person for logistics: case selection, spreadsheet, keeping to schedule. Nobody reviews cases they signed off themselves.
  2. Fix the sample. At least 40 real cases over 14 days, spread across weekdays and times of day. The largest part of it: unspectacular routine. Plus the six scenarios below, deliberately scheduled.
  3. Define the reference. Review always takes place against a source — for a conversation system, against the conversation itself (review on the same day, while memory still holds), for document AI, against the source document. Without a defined reference, you only measure whether the text sounds pleasing.
  4. Settle the legal side. Patient consent for every recorded conversation, a data processing agreement for the test period, a deletion concept for after the test — before the first case, together with the data protection officer and, where employee data are touched, early with the works council.
  5. Put error classes and stop criteria in writing — before the start. Criteria formulated only after the first annoyance are not criteria.

The six scenarios

Error rates of speech and document AI are strongly condition-dependent. In one review, the word error rate ranged from under 9 percent for controlled dictation to over 50 percent for conversational multi-speaker audio [1]. A stress test of a clinical transcription pipeline with 272 simulated encounters found, under babble noise at a signal-to-noise ratio of 5 decibels, a rise in word error rate from 16.5 to 54.7 percent — and a rise in clinically unsafe outputs from 13.6 to 91.5 percent [2]. These are simulation data from a preprint, but the direction is unambiguous: errors grow with the environment, and they grow nonlinearly. That is exactly why the difficult conditions are deliberately created in the trial.

  1. The family interview. Conduct several admission interviews with three or four people in the room. In benchmark data, the error rate of speaker separation nearly doubles when four voices have to be distinguished instead of two [3]. What to watch: does the daughter's account of symptoms end up attributed to the patient? Is "He doesn't take his tablets regularly" attributed to the right person — and marked as a relative's statement?
  2. Negations. Distribute deliberately negated statements across the test period: "no known allergies", "never smoked", "chest pain is denied", "Marcumar was discontinued". Count how many of them arrive in the draft as negations. Why this is the most expensive class of small errors is in Negation, laterality, dose (in German).
  3. Overlapping speech. Interject questions mid-sentence, let confirmations ("mhm", "exactly") run alongside. In meeting corpora, around 12 percent of speaking time overlaps; short confirmations are about 70 percent overlapped [4]. What to watch: what does the system make of overlapped passages — does it drop them, does it guess, does it blend two speakers?
  4. Background noise. Door open, corridor traffic, a running monitor. That is the difference between demo room and ward [2]. What to watch: does quality drop noticeably — and does the system say so, for instance via confidence indications?
  5. Paper quality (for document AI). Find the worst fax that actually arrived last month, a skewed scan, a handwritten note. What to watch: does the system mark uncertain extractions as uncertain, or does it deliver equally confident field values everywhere?
  6. The normal case. The majority of the 40 cases stays routine — otherwise you only measure the edges and cannot judge everyday performance.

Counting rules

Counting is per case, in five error classes:

  • Omission: a review-relevant item is in the source and missing from the draft. How to create a source list for this in advance is in Testing AI documentation for omissions (in German).
  • Invention: a statement in the draft has no counterpart in the source.
  • Person or speaker mix-up: a statement is attributed to the wrong person.
  • Flipped negation: negated becomes affirmed, or the reverse.
  • Detail error: dose, unit, side, date, numerical value.

Three rules make the count robust:

  1. Fix the denominator. What is reported is the share of cases with at least one error per class. Word error rates from data sheets are not comparable with this — differing error definitions alone make published rates mutually incomparable [5]. So never mix denominators.
  2. Two reviewers count independently; discrepancies are resolved by consensus, and the consensus result counts.
  3. Every error gets a relevance judgment: had it gone undetected, would it have changed treatment, billing or communication? Yes or no suffices.

Also record the review time per case. The cost of verification is part of the truth about the tool — a system that saves five minutes and needs seven minutes of review has a different value proposition than its brochure.

Stop criteria

In writing, before the start. Four criteria cover the critical cases:

  1. One invention with treatment relevance that the system presented without an uncertainty marking: pause the test, document the case to the vendor, continue only after a substantial answer.
  2. More than one flipped negation on medication or allergies: same procedure.
  3. One person mix-up that would have reached the record unmarked: same procedure.
  4. Zero errors found after 40 cases. The literature knows no error-free platform — in a comparison across five systems, the note error rate was 26.3 percent [6]. A spotless test result is therefore first of all a finding about the review procedure: time pressure, fatigue, too soft a reference. Check the checking before you trust the result.

After 14 days: three possible outcomes

  1. Introduce, with rules. The error profile is known; it becomes review rules for operations — which error classes the sign-off checks, which sample is re-measured monthly — and a date for the repeat measurement.
  2. Retest. Individual scenarios stood out, the rest held: a targeted second round only for the conspicuous conditions.
  3. End it. A stop criterion was triggered, and the vendor's answer does not hold.

In all three cases, one last step is worth it: give the vendor your count. The reaction is diagnostic. Whoever treats error figures from real operations as valuable material will behave the same way in later operations. Whoever explains them away, likewise.

If you want to train your review team beforehand: our CME-certified course on AI competence (3 CME points, Ärztekammer Hessen — the Hessian state medical chamber) teaches the foundations for exactly this review work. And in our weekly briefing Visite (German; English edition Grand Rounds is in preparation), we write regularly about how hospitals test and introduce AI systems.

Sources

  1. Übersichtsarbeit zur Spracherkennung in der klinischen Dokumentation. BMC Medical Informatics and Decision Making. 2025;25:236. https://pmc.ncbi.nlm.nih.gov/articles/PMC12220090/.
  2. Stresstest einer klinischen Transkriptionspipeline (272 simulierte Encounter). arXiv-Preprint, 2026. arXiv:2606.05909.
  3. Sortformer-Diarisierungsmodell, technische Dokumentation (CALLHOME-Benchmark, Fehlerrate nach Sprecherzahl). NVIDIA/Hugging Face. https://huggingface.co/nvidia/diar_sortformer_4spk-v1.
  4. Çetin Ö, Shriberg E. Analysis of overlaps in meetings by dialog factors, hot spots, speakers, and collection site: insights for automatic speech recognition. In: Proceedings of Interspeech 2006 — ICSLP; Pittsburgh, PA; 2006. doi:10.21437/Interspeech.2006-91.
  5. Topaz M, Peltonen LM, Zhang Z. Beyond human ears: navigating the uncharted risks of AI scribes in clinical practice. npj Digital Medicine. 2025;8(1):569. doi:10.1038/s41746-025-01895-6.
  6. Anderson TN, Mohan V, Dorr DA, Ratwani RM, Biro JM, Gold JA. Evaluating the Quality and Safety of Ambient Digital Scribe Platforms Using Simulated Ambulatory Encounters. Mayo Clinic Proceedings: Digital Health. 2025;3(4):100292. doi:10.1016/j.mcpdig.2025.100292.
#Clinical AI trial#AI test deployment hospital#AI test protocol#AI rollout hospital

Keep reading

An executive office at dusk with a packed appointment schedule on screen in the foreground, and a clinician pausing over a chart in a softly lit corridor behind the glass.
Reflections

The Jevons Paradox in Healthcare: Why Faster Doctors Are Not Better Doctors

When AI gives a clinician back ten minutes, the scheduling system tends to fill them with another patient. That instinct quietly converts every efficiency gain into more volume — and mistakes the bottleneck in medicine for time, when it was never time.

Dr. Sven JungmannCEO

This analysis comes from the people behind Visite.

Our weekly newsletter on AI in medicine. Every Friday, rigorously checked.

By signing up you agree to receive Grand Rounds by email. Unsubscribe anytime. More in our privacy policy.

Want to see this in your hospital?

30 minutes. Your questions. Our physician-founder shows you the platform personally.

Book a demo

No commitment. No sales pitch. Physician to physician.