J Imaging Inform Med. 2026 Sep 9. doi: 10.1007/s10278-026-02133-5. Online ahead of print.
ABSTRACT
Cardiovascular and cerebrovascular diseases are leading causes of death worldwide, motivating scalable, noninvasive screening approaches. Chest X-rays (CXRs) are widely available, but current deep learning methods rely on large labeled datasets that are costly to obtain in clinical practice. We propose a three-stage framework that combines DINOv2 self-supervised learning with vision transformers (ViTs) to detect stroke and heart failure from CXRs under limited labeled data. A ViT-Small/14 backbone is first initialized with DINOv2 weights pretrained on LVD-142M dataset and then further adapted by self-supervised pretraining on 5368 unlabeled CXRs from the UCSD dataset, with the bottom 8 transformer blocks frozen to preserve general visual representations during domain adaptation. Model evaluation is performed via frozen-backbone linear probing on 2000 labeled images per task from the NTUH-iMD database. Ablation experiments confirm that each pretraining stage contributes measurable gains under linear probing. Benchmarked against five SSL baselines spanning general domain and medical domain pretraining under an identical evaluation protocol, our stroke model achieves 89.60 ± 1.39% accuracy and 89.55 ± 1.40% F1-score, outperforming all baselines including RAD-DINO, which uses nearly four times the parameters. Our heart failure model achieves 92.52 ± 1.57% accuracy, surpassing GLoRIA, CheXzero, and RAD-DINO. We further show that negative class composition critically affects performance and that normal CXR proportion should be interpreted in light of its effect on task difficulty. Attention visualizations confirm that both models focus on anatomically meaningful cardiac and pulmonary regions, supporting the clinical interpretability of the proposed framework.
PMID:42717151 | DOI:10.1007/s10278-026-02133-5

