Development and validation of early warning scores predicting 24-hour mortality addressing intercurrent medical interventions: a multi-centre retrospective cohort study of 2,08 million hospital encounters
Thursday, October 1, 2026
BackgroundThe National Early Warning Score (NEWS) is widely used to predict patient deterioration, but retrospective validations face intervention bias and emphasise discrimination over clinical utility. We aimed to compare NEWSs clinical net benefit against simplified scoring rules and machine learning, to determine whether simpler approaches suffice or if complex models offer meaningful advantages, and to evaluate performance heterogeneity across patient subgroups. MethodsWe included fifteen Danish hospitals with over 2{middle dot}08 million hospital encounters representing 825200 unique patients over five years (2018 to 2023). We compared NEWS against both simpler and more complex approaches for predicting 24h mortality: Simplified NEWS (NEWS without blood pressure and temperature), DEWS (Simplified NEWS with age and sex), and a model based on eXtreme Gradient Boosting (XGB-EWS) incorporating vital signs, demographics, laboratory markers, plus medical history embeddings extracted using sentence transformers. We used propensity score weighting to mitigate intervention bias and evaluated performance using net benefit, Area Under the Receiver Operating Characteristic Curve (AUC), and calibration. FindingsDecision curve analysis in the overall population showed maximum net benefit differences of 1{middle dot}9 additional correct mortality identifications per 10000 patients between XGB-EWS and NEWS, and 1{middle dot}5 per 10000 between NEWS and Simplified NEWS. However, stratified analyses revealed substantial heterogeneity: in the high-risk group (initial NEWS [≥]7), XGB-EWS provided a maximum gain of 78{middle dot}1 correct identifications per 10000 patients compared to NEWS, while Simplified NEWS resulted in a maximum performance loss of 60{middle dot}2 per 10000. Conversely, in low-risk groups, differences were minimal. InterpretationModel complexity produced lim...
Daily healthcare AI brief
Get the free healthcare AI briefing
One concise briefing built from reporting, research, policy, blogs, videos, and full podcast transcripts.
Free. Sign in to subscribe. Unsubscribe anytime. Already subscribed? Sign in to save and personalize.