Back to OpenRounds
PreprintmedRxivPreprint

Reasoning Before Disposition: A Model-Agnostic Cannot-Miss Discipline for Quiet Emergencies and the Case for Deterministic Enforcement

Tuesday, September 8, 2026

Referenced in Daily Briefing

Large language models now match clinicians on medical-knowledge benchmarks. Whether they can safely triage a patient message is a separate question, and the highest-consequence failure in triage is the emergency that never gets escalated. We tested two competing explanations for that failure. The first is sycophancy: the model defers to a patient who downplays a danger sign. The second is atypicality: the model reads an early, mild, or atypical presentation of a time-critical disease as benign. To isolate the effect of a model-agnostic cannot-miss discipline, we performed a prompt-level ablation on a fixed frontier model (Claude Opus 4.8), then replicated the same instruction-level intervention across eight models from two families (Claude Opus 4.8, Sonnet 5, Haiku 4.5, and Fable 5; Gemini 3.1 Pro-Preview, 3.6 Flash, 3.5 Flash-Lite, and 3.1 Flash-Lite) under a documented harness and assembled the full apparatus as a model-agnostic evaluation suite. Minimization turned out to be harmless when danger is overt. Across 40 unambiguous emergencies rendered in five framings on Opus 4.8 (neutral, mild minimization, strong denial, third-party reassurance, and a benign distractor), sensitivity stayed at or above 97.5% in every arm and was flat across framings, and seven of the eight models held 97.5% to 100% on the neutral set (the eighth completed 13 of 40 stems, all escalated). Atypicality is where the hazard lives. Across 76 subtle or atypical emergencies, each adjudicated by a blinded, LLM-simulated three-physician panel and each mapped to a published guideline that names its cannot-miss diagnosis, the base model failed to escalate 13.2% (95% CI 7.3 to 22.6) to the top acuity tier; the governance prompt cut this to 3.9% (1.4 to 11.0; paired exact McNemar 7 to 0, p = 0.016; relative risk 0.30; number needed to treat 11). The effect sits in the harder half of the cohort (36 o...

Read Source ↗

OpenRounds Source Analysis

Sign in for an evidence-focused interpretation, plus Saved articles and your personalized feed.

Sign InBack to OpenRounds