Comparative Performance of AI Models and Clinicians in Evidence-Based Cardiovascular Disease Management for People Living With HIV: Comparative Study

Scritto il 10/08/2026
da Tianqi Kong

J Med Internet Res. 2026 Aug 10;28:e89858. doi: 10.2196/89858.

ABSTRACT

BACKGROUND: Although widespread antiretroviral therapy has extended the life expectancy of people living with HIV, cardiovascular disease (CVD) has emerged as a primary comorbidity. Persistent cross-specialty knowledge gaps in routine clinical practice lead to suboptimal adherence to guidelines. Integrated, evidence-based tools are urgently needed to overcome these interdisciplinary barriers. While large language models (LLMs) have demonstrated significant capabilities in medicine, no systematic evaluation has assessed their ability to facilitate multidisciplinary CVD management for people living with HIV.

OBJECTIVE: This study compared the performance of 4 mainstream AI models (DeepSeek-V3, DeepSeek-R1, ChatGPT-4o, and ChatGPT-o4-mini) against 12 human clinicians (8 infectious disease specialists and 4 cardiologists) in addressing guideline-based CVD management tasks for people living with HIV.

METHODS: Based on 4 authoritative domestic and international HIV/CVD guidelines, a structured 25-question assessment was developed via 2 rounds of Delphi consultation. Standard reference answers and an evaluation framework were finalized through expert consensus. LLM responses were generated using standardized prompts. Clinicians answered identical questions via one-on-one structured interviews, transcribed verbatim. Six multidisciplinary experts independently rated all responses across 4 dimensions-accuracy, completeness, readability, and reliability-using a 4-point ordinal scale (1=poor to 4=excellent). Cumulative link mixed models analyzed intergroup differences.

RESULTS: All AI models achieved significantly higher scores than clinicians across all dimensions (P<.001). The AI group's mean scores ranged from 3.44 to 3.68 (median 4, IQR 3.0-4.0; coefficient of variation=0.145-0.178). Conversely, clinicians' scores were lower (mean 1.78-2.05; median 2, IQR 1.0-3.0; coefficient of variation=0.428-0.473) with marked dispersion. DeepSeek-R1 delivered the optimal performance, significantly outperforming the other 3 models (all P<.001). Specialty-stratified analysis revealed no significant overall score difference between cardiologists and infectious disease specialists (odds ratio 0.92, 95% CI 0.84-1.01; P=.09). However, dimension-specific analysis indicated that cardiologists scored higher in accuracy (odds ratio 0.81, 95% CI 0.67-0.97; P=.03). Domain-specific divergence was evident: cardiologists outperformed infectious disease specialists in CVD risk assessment (2.26 vs 1.83), whereas infectious disease specialists led in drug adverse effect evaluation (2.23 vs 1.65).

CONCLUSIONS: In this structured question-and-answer study, LLMs outperformed human clinicians across all metrics for HIV-associated CVD management, with DeepSeek-R1 achieving superior composite scores. These findings validate DeepSeek-R1's potential as a cross-disciplinary decision-support tool capable of integrating complex clinical knowledge, mitigating specialty gaps, and enhancing information precision. Integrating AI systems into multidisciplinary workflows, complemented by targeted clinical training, may optimize the management of complex comorbidities in people living with HIV.

PMID:42574667 | DOI:10.2196/89858