Guilin Multi-ML
204 SAP / 1,085 non-SAP. Approximately 60 predictors. Raw Diagnostic Result uses 0=SAP and 1=non-SAP; PenuX explicitly normalizes to 1=SAP.
The program targets roughly 3,000 nominal acute-pancreatitis records while keeping cohort identity, source licensing, outcome definitions and possible overlap explicit.
204 SAP / 1,085 non-SAP. Approximately 60 predictors. Raw Diagnostic Result uses 0=SAP and 1=non-SAP; PenuX explicitly normalizes to 1=SAP.
137 SAP / 585 non-SAP, about 107 predictors. Related institutional cohort and not treated as fully independent external validation.
200 development + 60 validation patients with laboratory, demographic and clinical variables and Atlanta-based severity labels. Institutionally independent from Guilin.
Multi-center US ICU source with time-stamped labs and diagnoses. Credentialed data are never redistributed; PenuX must derive its own Atlanta-compatible severity cohort.
Counts from the two Guilin datasets cannot simply be assumed to represent disjoint patients because the institution and calendar periods overlap. The eICU figure is a published candidate AP cohort rather than a ready-made PenuX SAP outcome cohort.
Map aliases such as creatinine/Cr, glucose/Glu, albumin/ALB and laboratory unit variants to a canonical schema.
Normalize each source to 0=non-SAP and 1=SAP only after verifying the source coding and outcome definition.
Retain only predictors available at the intended admission/ward time point; do not leak future organ-failure measurements.