Digit Health. 2026 Jul 17;12:20552076261456893. doi: 10.1177/20552076261456893. eCollection 2026 Jan-Dec.
ABSTRACT
BACKGROUND: Large language models are increasingly applied in medical education and clinical decision support. However, comparative evaluations of their performance on pediatric content-spanning multiple subspecialties and question complexities-remain limited. This study sought to assess the performance of two advanced large language models, DeepSeek-R1 and GPT-4o, using a comprehensive set of Pediatrician Licensing Exam practice questions.
METHODS: We administered 280 expert-validated pediatric questions covering 11 subspecialties and three levels of clinical complexity (A1: direct, A2: simple cases, A3: complex cases) to DeepSeek-R1 and GPT-4o. Each model was tested twice, with a 4-week interval. Performance metrics, including per-run accuracy, consistent accuracy (correct in both runs), aggregate accuracy (correct in either run), and inter-run consistency, were compared using chi-square tests and Bonferroni correction.
RESULTS: DeepSeek-R1 outperformed GPT-4o in both runs (1st run: 90.7% vs. 81.8%, P=0.003; 2nd run: 86.8% vs. 79.3%, P = 0.02). It also showed higher aggregate accuracy (92.5% vs. 85.7%, P=0.01) and consistent accuracy (85.0% vs. 75.4%, P = 0.01). Performance declined with increasing question complexity for both models, with DeepSeek-R1 demonstrating numerically higher accuracy across all question types. For consistent accuracy, DeepSeek-R1 achieved 89.2% vs. 79.2% on A1 questions (P=0.05), 83.8% vs. 71.6% on A2 questions (P=0.11), and 80.2% vs. 73.3% on A3 questions (P=0.36). At the subspecialty level, DeepSeek-R1 showed more uniform performance, achieving >90% consistent accuracy in four subspecialties: Neonatology (94.7%), Infectious Diseases (93.8%), Nephrology (91.9%), and Cardiovascular diseases (90%). In comparison, GPT-4o reached >90% consistent accuracy in three subspecialties-Genetics (100%), Infectious Diseases (100%), and Cardiovascular diseases (90%)-with greater variability across other domains.
CONCLUSION: DeepSeek-R1 demonstrated higher accuracy than GPT-4o on pediatric questions. The observed variability across different runs warrants attention, particularly in contexts that require highly consistent and reproducible performance.
PMID:42472269 | PMC:PMC13379656 | DOI:10.1177/20552076261456893