1. Development
Fit preprocessing and tune XGBoost only inside development folds. Class weighting and feature selection are development operations.
The research design separates development, threshold selection and final evaluation. A high AUROC alone is insufficient for an early-warning claim.
Fit preprocessing and tune XGBoost only inside development folds. Class weighting and feature selection are development operations.
Generate out-of-fold probabilities to compare configurations, assess calibration and select the high-sensitivity operating threshold.
Freeze model identity and threshold before evaluation on held-out or external cohorts.
The final report must give specificity, PPV, NPV and alert burden at this same locked threshold—not at a retrospectively chosen test-set threshold.
| Domain | Measures | Why it matters |
|---|---|---|
| Discrimination | AUROC, AUPRC | AUPRC is especially informative when SAP is the minority class. |
| Operating point | Sensitivity, specificity, PPV, NPV, Fβ | Shows the actual cost of targeting very high sensitivity. |
| Calibration | Brier score, slope, intercept, reliability curve | A ranked model can still give badly biased absolute probabilities. |
| Uncertainty | Bootstrap confidence intervals | Prevents over-reading point estimates from modest cohorts. |
| Utility | Decision-curve analysis, alert burden | Explores whether a threshold may add net benefit, without implying clinical approval. |
The preferred design is leave-one-cohort-out or explicitly external validation. Cohort identity is used for splitting and reporting, not as a shortcut predictor.
Performance should be shown for each hospital/cohort before any pooled summary. The two Guilin sources require particular caution because they may contain overlapping patients.
For time-stamped EHR cohorts, predictions can be evaluated at admission, 6 h, 12 h and 24 h. Events already present at time \(t\) must not receive “early warning” credit.