Front Neurol. 2026 Aug 14;17:1847449. doi: 10.3389/fneur.2026.1847449. eCollection 2026.
ABSTRACT
BACKGROUND: Deep vein thrombosis (DVT) is a common complication of acute ischemic stroke (AIS) and may worsen clinical outcomes, yet reliable tools for early risk stratification remain limited. We aimed to develop and internally evaluate machine learning models for predicting in-hospital DVT in patients with AIS.
METHODS: We conducted a secondary analysis of a publicly available multicenter retrospective dataset including 21,459 patients with AIS. The primary outcome was imaging-confirmed in-hospital DVT. Participants were stratified according to DVT status and randomly divided into a training set (70%) and a held-out test set (30%). Feature selection was performed using least absolute shrinkage and selection operator regression, the Boruta algorithm, variance inflation factor assessment, and clinical judgment. Eight machine learning algorithms were trained using a 19-variable full predictor set and an 8-variable simplified predictor set. Hyperparameters were optimized using repeated 5-fold cross-validation with 2 repeats. Model performance was evaluated using the area under the receiver operating characteristic curve (AUC), area under the precision-recall curve (AUPRC), Brier score, calibration, decision curve analysis, and additional classification metrics. Sensitivity analyses excluded D-dimer and used within-fold synthetic minority oversampling.
RESULTS: Among 21,459 patients, 1,324 (6.17%) developed in-hospital DVT. In the full predictor-set analysis, RANGER achieved the highest AUC in the held-out test set (0.976). Among models using the simplified predictor set, XGBoost achieved the highest AUC (0.917) and sensitivity (0.852), whereas SVM demonstrated the most favorable overall performance profile, with the highest AUPRC (0.605), lowest Brier score (0.038), highest positive predictive value (0.440), and highest F1-score (0.549). D-dimer was the most influential predictor. Model performance was attenuated after exclusion of D-dimer, and the SMOTE analysis showed slightly lower discrimination and greater calibration discrepancies. All 8 simplified predictor-set models were incorporated into an online prediction platform.
CONCLUSION: Machine learning models demonstrated favorable performance for predicting in-hospital DVT after AIS. Among models using the simplified 8-variable predictor set, SVM showed the most favorable overall performance, whereas XGBoost prioritized sensitivity. Independent external validation and prospective clinical-impact assessment are required before routine clinical implementation.
PMID:42666172 | PMC:PMC13521839 | DOI:10.3389/fneur.2026.1847449