Sci Rep. 2026 Jul 29. doi: 10.1038/s41598-026-62750-6. Online ahead of print.
ABSTRACT
Cardiovascular diseases (CVDs) remain a leading cause of mortality worldwide, motivating reproducible benchmark studies on accurate classification methods with calibrated probability estimates. In this paper, we propose a deep learning framework based on a Variational Recurrent Autoencoder (VRAE) with uncertainty estimation for CVD classification from static tabular clinical records represented as synthetic noise-augmented pseudo-sequences. The two datasets used in this study are static tabular datasets rather than real longitudinal clinical time-series. Therefore, the constructed pseudo-sequence dimension is used for denoising latent representation learning under simulated measurement perturbation and missingness, not for modeling observed clinical temporal dependencies, patient trajectories, disease progression, or treatment dynamics. The model integrates Gated Recurrent Unit (GRU)-based encoding, variational latent sampling, and Monte Carlo (MC) dropout to learn robust latent representations and quantify predictive uncertainty. We evaluate the model on two publicly available datasets: the Heart Failure Prediction dataset (918 samples) and the Cardiovascular Disease dataset (70,000 samples). The model achieves an accuracy of 95.8% and F1-score of 95.7% on the heart failure dataset, and an accuracy of 96.1% with F1-score of 96.0% on the larger cardiovascular dataset, outperforming traditional classifiers, calibrated tabular baselines, and deep learning baselines including regularized logistic regression, calibrated XGBoost, calibrated LightGBM, calibrated CatBoost, Long Short-Term Memory (LSTM), and GRU-Attention. The proposed VRAE model also demonstrates robustness under synthetic noise and provides improved calibration, as indicated by the lowest Brier scores (0.061 and 0.059) across both datasets. Additional calibration, selective prediction, and threshold-based utility analyses support the reliability of the benchmark results. These findings support uncertainty-aware representation learning for benchmark-level static tabular CVD classification, while independent validation on real-world longitudinal electronic health record cohorts is required before claims about prospective risk prediction, early intervention, clinical time-series prediction, or clinical deployment can be made.
PMID:42477085 | DOI:10.1038/s41598-026-62750-6