Int J Med Inform. 2026 Jul 24;220:106635. doi: 10.1016/j.ijmedinf.2026.106635. Online ahead of print.
ABSTRACT
BACKGROUND: Acute coronary syndrome (ACS) causes substantial post-discharge mortality. Ambient air pollution and meteorological conditions are associated with recurrent cardiovascular events, but existing clinical risk scores rely only on static admission parameters without incorporating post-discharge environmental exposures. Most machine learning (ML) models for ACS prognosis are single-center, single-horizon, fail to integrate both meteorological and air-quality variables, and lack external validation or interpretability.
METHODS: We developed an interpretable multi-horizon ML framework for 3-month, 6-month, and 1-year post-discharge mortality using 24,120 ACS patients from 28 institutions in the Tianjin Chest Pain Center registry, split into training (n = 15,436), internal validation (n = 3,860), and single-region, multi-institution external validation (n = 4,824) cohorts. Seventeen predictors (3 demographic/clinical, 14 environmental) were evaluated across 5 algorithms. Discrimination was assessed by AUC (95% CI, DeLong method), calibration by Brier scores and calibration plots, and survival differences by log-rank test. Missing data were handled via multiple imputation, with SHAP analysis quantifying predictor contributions.
RESULTS: CatBoost demonstrated the most stable cross-cohort performance, with 1-year AUCs of 0.625 (0.596-0.654) and 0.629 (0.602-0.654) in internal and external validation, significantly outperforming random forest and XGBoost (P < 0.05). Top predictors included winter minimum relative humidity, length of hospital stay, summer mean temperature, sex, and annual PM Predefined risk thresholds yielded significantly separated Kaplan-Meier curves (log-rank P < 0.001). Given missing key admission-severity indicators (Killip class, cardiac biomarkers, ECG findings), SHAP-based rankings are conditional on the available feature set and may overstate environmental exposures' relative contribution.
CONCLUSIONS: With moderate discrimination (AUC 0.62-0.65) and single-region design, this multi-horizon ML framework integrating meteorological and air-quality variables was validated via single-region, multi-institution external validation. It is intended as a low-cost triage signal to prioritize follow-up, not a stand-alone decision tool, and does not replace clinical judgment. Cross-regional validation is required.
PMID:42508148 | DOI:10.1016/j.ijmedinf.2026.106635