Multi-cohort expansion

Approximately 3,000 AP records across four cohorts.

PenuX-AP-Severity is being expanded from a single-dataset experiment into a cohort-aware validation framework. The objective is not simply to increase sample size, but to test whether an early XGBoost signal is stable across institutions, case mixes and data-collection systems.

Nominal N = 3,0174 source cohortsLabs + diagnosesXGBoostAtlanta-compatible SAP outcomeTarget sensitivity ≥98%
Important: 3,017 is a planning total of source records, not yet a verified count of unique eligible patients. The two Guilin cohorts may contain overlapping individuals, and eICU will lose records after eligibility/missingness/outcome-label checks.

Core cohort composition

public / repository
1,289
Guilin Multi-ML

204 SAP; ~60 predictors. Main tabular laboratory benchmark.

public / repository
722
Guilin LNN

137 SAP; ~107 predictors. Related Guilin cohort with a richer feature set.

public source
260
Hefei / Han et al.

200 development + 60 validation records with labs, clinical variables and SAP/NSAP labels.

credentialed public data
746*
eICU-CRD AP

Published candidate AP cohort from a 208-hospital US ICU database; final PenuX N must be re-derived.

Why the datasets must not simply be concatenated

CohortSettingStrengthMain limitationRecommended role
Guilin Multi-MLSingle institution, ChinaLargest directly usable labeled tabular cohortSingle-center; source-specific measurement patternsPrimary development / internal validation
Guilin LNNSame institution, overlapping calendar periodMore predictors and additional recordsPossible patient overlap with Multi-ML cannot be excluded after de-identificationSensitivity analysis; not an independent external cohort
Hefei / OSFDifferent Chinese institutionInstitutionally independent source with labs and severity labelsSmaller sampleExternal / transportability validation
eICU-CRD208 US hospitals, ICU-enrichedStrong geographic and institutional shift; timestamped labs and diagnosesCredentialed access; ICU spectrum differs from general AP; Atlanta-compatible SAP label must be derivedMulti-center stress test and time-aware validation

Validation architecture

Source-specific QC
units, ranges, missingness
Common lab schema
harmonized predictors
Development CV
XGBoost tuning
Locked model
98% sensitivity threshold
Hefei test
external site
+
eICU test
US multicenter shift
Per-cohort report
AUROC · AUPRC · calibration
Pooled summary
only after heterogeneity

Common laboratory feature set

A cross-cohort model should prioritize measurements that can be mapped reliably across datasets. Candidate common variables include WBC, hemoglobin/hematocrit, platelets, BUN/urea, creatinine, glucose, calcium, sodium, potassium, albumin, bilirubin, AST/ALT, LDH, coagulation tests and selected vital signs. Features that exist only in one cohort can be evaluated in cohort-specific secondary models but should not define the primary transportability model.

Leakage rule: predictors must be available before the prediction time. Persistent organ failure or later interventions used to define the future SAP outcome cannot re-enter the model as predictors.

Outcome compatibility

The Chinese labeled cohorts already provide SAP/non-SAP outcomes tied to acute-pancreatitis severity definitions. eICU does not provide a ready-made Revised Atlanta SAP column. For the primary severity analysis, PenuX should derive a persistent-organ-failure endpoint (>48 h) from time-stamped organ-failure measurements where feasible. ICU admission or mortality alone must not be relabeled as SAP.

For eICU, mortality can be reported as a separate secondary outcome. This preserves scientific interpretability and prevents incompatible labels from being pooled.

Additional auxiliary dataset

auxiliary · not counted in core N
538
MIMIC-IV-Ext pancreatitis cases

MIMIC-IV-Ext Clinical Decision Making contains 538 pancreatitis cases and extensive laboratory results. It is useful for laboratory-schema mapping, diagnosis tasks and representation/domain experiments, but it is not counted in the ~3,017 SAP pool until an outcome compatible with the PenuX severity definition is established.

PhysioNet source

What success would mean

MetricDevelopment goalWhy it matters
Sensitivity≥98% at locked watch thresholdMinimize missed future severe cases in the research early-warning setting.
SpecificityReport, do not forceShows the alert burden created by high sensitivity.
AUROCDiscrimination across cohortsMeasures rank discrimination but does not define the alert threshold.
AUPRCPrimary model-selection metricMore informative when SAP is the minority class.
CalibrationSlope/intercept + Brier + reliabilityRequired before interpreting XGBoost output as absolute risk.
External stabilityNo large collapse on Hefei/eICUMore important than a marginal gain on the original development cohort.