Multi-cohort evidence base

Datasets and provenance

The program targets roughly 3,000 nominal acute-pancreatitis records while keeping cohort identity, source licensing, outcome definitions and possible overlap explicit.

3,017 is a planning total, not a confirmed unique-patient analytic N. Final N requires eligibility filtering, outcome harmonization and overlap checks.
1,289

Guilin Multi-ML

204 SAP / 1,085 non-SAP. Approximately 60 predictors. Raw Diagnostic Result uses 0=SAP and 1=non-SAP; PenuX explicitly normalizes to 1=SAP.

722

Guilin LNN

137 SAP / 585 non-SAP, about 107 predictors. Related institutional cohort and not treated as fully independent external validation.

260

Hefei / Han et al.

200 development + 60 validation patients with laboratory, demographic and clinical variables and Atlanta-based severity labels. Institutionally independent from Guilin.

746

eICU AP candidate cohort

Multi-center US ICU source with time-stamped labs and diagnoses. Credentialed data are never redistributed; PenuX must derive its own Atlanta-compatible severity cohort.

Nominal planning total

1,289 + 722 + 260 + 746 = 3,017 source records

Counts from the two Guilin datasets cannot simply be assumed to represent disjoint patients because the institution and calendar periods overlap. The eICU figure is a published candidate AP cohort rather than a ready-made PenuX SAP outcome cohort.

Open the detailed multi-cohort plan →

Cross-cohort harmonization

Feature names

Map aliases such as creatinine/Cr, glucose/Glu, albumin/ALB and laboratory unit variants to a canonical schema.

Outcome labels

Normalize each source to 0=non-SAP and 1=SAP only after verifying the source coding and outcome definition.

Prediction time

Retain only predictors available at the intended admission/ward time point; do not leak future organ-failure measurements.