Evaluation of Artificial Intelligence (AI) for Myocardial Infarction Health Information: Expert Accuracy Assessment and User Trust Survey

Scritto il 28/08/2026
da Eman Elsheikh

Clin Ter. 2026 Sep-Oct;177(5):1216-1233. doi: 10.7417/CT.2026.2124.

ABSTRACT

BACKGROUND: Artificial intelligence chatbots, particularly ChatGPT, have emerged as increasingly popular sources of health information for the general public. However, concerns persist regarding the accuracy, safety, and appropriateness of AI-generated medical advice, especially for life-threatening conditions such as myocardial infarction.

OBJECTIVES: This study aimed to evaluate the accuracy, completeness, and safety of ChatGPT-generated responses and information to public questions about heart attacks, assess user perceptions and trust, and compare AI-generated content with expert-validated sources.

METHODS: A cross-sectional study was conducted between January and March 2026. A total of 73 commonly asked heart attack-related questions were submitted to ChatGPT (GPT-4), and responses were independently evaluated by two expert reviewers and one external reviewer using standardized rubric assessing accuracy, completeness, clarity, safety, and tone. Inter-rater reliability was assessed using intraclass correlation coefficients. In parallel, a survey of 352 participants evaluated user perceptions, trust, and behavioral use of ChatGPT for medical information. Statistical analyses included descriptive statistics, group comparisons, correlation analyses, and multivariable regression models to identify predictors of trust and satisfaction.

RESULTS: Expert reviewers assigned high ratings for accuracy (mean: 4.62 ± 0.62 and 4.48 ± 0.63) and safety (mean: 4.78 ± 0.51 and 4.05 ± 0.50), with substantial inter-rater agreement (ICC = 0.788). In contrast, the external reviewer documented considerably lower completeness ratings (2.59 ± 0.57), revealing deficiencies in information coverage. Remarkably, 40.6% of respondents indicated they had followed ChatGPT's medical recommendations without seeking professional consultation, while 17.6% reported depending on the chatbot during urgent medical situations. Levels of user trust and satisfaction were moderate, showing significant correlations with perceived accuracy, usage frequency, and healthcare-related educational background.

CONCLUSIONS: Expert assessments confirmed that ChatGPT generally provides accurate and safe cardiovascular health information; nevertheless, considerable inter-reviewer variability-especially the markedly lower completeness scores from the external evaluator-reveals significant gaps in content depth. These results indicate that ChatGPT functions more appropriately as an adjunctive educational resource rather than a replacement for professional medical consultatio n, emphasizing the necessity for enhanced safety mechanisms, clearer user instructions, and standardized content protocols when addressing high-stakes health topics. ChatGPT should be considered an adjunctive resource rather than a substitute for urgent professional medical assessment in suspected heart attack cases.

PMID:42664144 | DOI:10.7417/CT.2026.2124