Lancet Digit Health. 2026 Aug 28:101031. doi: 10.1016/j.landig.2026.101031. Online ahead of print.
ABSTRACT
BACKGROUND: RETFound, a self-supervised retina-specific foundation model, has shown potential in downstream tasks, but its performance in comparison with that of traditional deep-learning models remains unclear. We aimed to evaluate RETFound against three supervised deep-learning models (ResNet50, ViT-Base, and SwinV2) pretrained on images from ImageNet for the detection of several ocular (disease-related visual impairment, visually significant cataract, glaucoma, and diabetic retinopathy) and systemic (diabetes, hypertension, and chronic kidney disease) diseases.
METHODS: In this retrospective comparative study, all models were fine-tuned on different proportions of the training dataset (100%, 50%, and 20%) and on smaller datasets of fixed sizes (100, 200, and 400 images; 250 and 500 images for diabetic retinopathy) for all tasks. The fine-tuned models were tested on internal datasets from the Singapore Epidemiology of Eye Disease (SEED) study (1842-10 474 images) for all diseases except diabetic retinopathy, for which the Asia Pacific Tele-Ophthalmology Society 2019 dataset (1100 images) was used. Each model was also evaluated on external datasets, comprising population-based datasets (from the Beijing Eye Study, the Central India Eye and Medical Study, the Singapore Prospective Study, and the UK Biobank) and open-source datasets (Ocular Disease Recognition-5K [ODIR-5K], PAPILA, Glaucoma Grading from Multi-Modality Images [GAMMA], Indian Diabetic Retinopathy Image Dataset, and Methods to Evaluate Segmentation and Indexing Techniques in the Field of Retinal Ophthalmology-2). The performance of the models was assessed using 2000 pairwise bootstrap iterations of the area under the receiver operating characteristic curve (AUC) and compared using a two-tailed test with Bonferroni correction, with statistical significance defined as p<0·017 to account for the three pairwise comparisons between models.
FINDINGS: In internal testing, similar performance was observed for traditional models (AUC 0·914 [95% CI 0·899-0·928] to 0·965 [0·952-0·976]) and RETFound (0·938 [0·920-0·954] to 0·966 [0·951-0·978]) in the detection of ocular disease after fine-tuning on full datasets. With smaller datasets, the performance of all models was similar, except in the case of diabetic retinopathy (≤100 images per class) and glaucoma (≤400 images), for which ResNet50 was inferior to RETFound (all p≤0·0001), although the performance of SwinV2 remained similar to that of RETFound. Similar patterns were observed with external test sets. As an example for glaucoma detection, after fine-tuning on 400 images, RETFound had an AUC of 0·908 (95% CI 0·888-0·928) when tested on the internal SEED dataset and, for the external datasets, AUCs of 0·835 (0·816-0·855) when tested on ODIR-5K, 0·779 (0·731-0·826) on PAPILA, and 0·990 (0·974-1·000) on GAMMA, performing significantly better than ResNet50 (SEED p=0·0001; ODIR-5k p<0·0001; PAPILA p=0·0032; and GAMMA p=0·0010). For systemic diseases, RETFound consistently outperformed traditional deep-learning models in internal testing when fine-tuned on smaller datasets (≤400 images): for example, for the detection of hypertension when fine-tuned on 100 images, the AUC for RETFound in the SEED dataset was 0·705 (0·687-0·723), compared with 0·634 (0·615-0·654) for ResNet50 and 0·648 (0·628-0·667) for SwinV2 (both p<0·0001) and 0·657 (0·637-0·676) for ViT-Base (p=0·0003). When tested in the external UKBB dataset, RETFound achieved an AUC of 0·622 (95% CI 0·618-0·625) for the detection of hypertension when fine-tuned on 400 images, 0·599 (0·595-0·602) on 200 images, and 0·599 (0·595-0·603) on 100 images, significantly outperforming ResNet50 and SwinV2 when fine-tuned on 400, 200, and 100 images (all p≤0·0001) and ViT-Base when fine-tuned on 400 images (p<0·0001) and on 200 images (p=0·0002).
INTERPRETATION: Under this specific study design, the performance of traditional deep-learning models is similar to that of RETFound for ocular disease detection when fine-tuned on large datasets. By contrast, RETFound shows an advantage in the detection of systemic disease when fine-tuned on smaller datasets. These findings offer insights into the respective merits and limitations of traditional models and foundation models. Future benchmarking on broader datasets is warranted.
FUNDING: Agency for Science, Technology and Research (A∗STAR).
PMID:42665469 | DOI:10.1016/j.landig.2026.101031

