Lancet Digit Health. 2026 Sep 1:101092. doi: 10.1016/j.landig.2026.101092. Online ahead of print.
ABSTRACT
BACKGROUND: Automated electrocardiogram (ECG) assessment tools to assist clinicians in diagnosis have improved substantially over the past decade; however, the full diagnostic and predictive potential of ECGs is limited when using traditional machine learning approaches because of an over-reliance on task-specific labels. We aimed to develop a large-scale ECG foundation model, ECG-CLIP, with improved generalisability and clinical applicability for ECG analysis.
METHODS: This study used data from the Scripps Health GE MUSE system, a large-scale, retrospective dataset containing more than 1·7 million ECGs and paired clinician-overread ECG reports from 542 288 patients, collected in-clinic between Jan 15, 2008, and Jan 15, 2019. ECG-CLIP was pretrained in a two-stage pretraining framework consisting of masked ECG reconstruction and contrastive learning combining ECG data with textual annotations. MIMIC-IV, an independent, external dataset with more than 800 000 ECGs, was used to evaluate ECG-CLIP in cardiovascular disease detection (acute myocardial infarction, cardiac amyloidosis, and hypertrophic cardiomyopathy), cardiovascular disease prediction (atrial fibrillation from normal sinus rhythm), and adverse health outcome prediction (30-day emergency department and post-surgical mortality and 3-year onset of chronic kidney disease and type 2 diabetes) tasks. We compared the performance of ECG-CLIP against that of two supervised baseline models and one non-ECG foundation model, as well as three ECG signal-only foundation models: ECGFounder, DeepECG-SL, and DeepECG-SSL. Performance was evaluated using a five-fold cross-validation framework and reported as the area under the receiver operating characteristic curve (AUC). We assessed label efficiency by varying the number of positive labels in the training set of each downstream task.
FINDINGS: ECG-CLIP consistently showed superior performance across all detection and prediction tasks compared with the supervised baseline and non-ECG foundation models, as well as a high label efficiency, reaching the same AUC in cardiovascular disease detection tasks as the best-performing comparator model with an average of 90·8% less training data. Compared with ECG signal-only foundation models, ECG-CLIP showed superior performance in low-label settings. When only ten positive labels were present in the dataset, ECG-CLIP significantly outperformed the runner-up model on tasks including acute myocardial infarction detection (AUC 0·910 [95% CI 0·903-0·916] vs 0·884 [0·876-0·892]), cardiac amyloidosis detection (0·790 [0·767-0·812] vs 0·777 [0·754-0·799]), hypertrophic cardiomyopathy detection (0·772 [0·751-0·793] vs 0·754 [0·733-0·774]), and atrial fibrillation prediction (0·777 [0·773-0·781] vs 0·765 [0·761-0·769]). The model has good performance in the detection of anterior and inferior acute myocardial infarction when using single ECG leads, outperforming supervised baseline models when using leads typically considered less informative.
INTERPRETATION: ECG-CLIP is a generalisable foundation model that can learn clinically meaningful representations from ECGs and paired reports for diagnostic and prediction tasks. This approach enables accurate, data-efficient, and interpretable risk predictions across diverse clinical tasks, supporting scalable deployment in real-world and resource-limited settings.
FUNDING: The VoLo Foundation and the National Center for Advancing Translational Sciences at the National Institutes of Health.
PMID:42680677 | DOI:10.1016/j.landig.2026.101092

