Evaluation of a National Health Service Machine-Learning Model for Hypertension Case-Finding: Retrospective Cohort Study

Scritto il 15/09/2026
da Gloria Ihenetu

J Med Internet Res. 2026 Sep 15;28:e87084. doi: 10.2196/87084.

ABSTRACT

BACKGROUND: Hypertension is a leading preventable cause of cardiovascular disease, yet a substantial proportion of adults remain undiagnosed, limiting opportunities for early intervention. A predictive model was commissioned by the North West London (NWL) Integrated Care Board to identify undiagnosed hypertension. The model was developed using health records from the Whole Systems Integrated Care (WSIC) database.

OBJECTIVE: We aimed to independently evaluate the predictive performance of the model as it would be encountered in deployment, how performance varied by demographic characteristics, and practical utility.

METHODS: To evaluate the predictive model, we conducted a retrospective cohort study of 1,802,920 individuals aged 16 years or older, registered with a general practice in NWL, and with no prior diagnosis of hypertension from May 2023 to May 2024. We assessed the model's predictions against recorded hypertension status using medical diagnoses and blood pressure records. Logistic regression models were used to assess the sensitivity and specificity of the model's predictions by sociodemographic groups. We also compared the model's performance against a more interpretable regression approach.

RESULTS: The model yielded an overall sensitivity of 62.7% (95% CI 62.5-62.8) and specificity of 60.7% (95% CI 60.5-60.8). Positive predictive value ranged from 31.5% (95% CI 31.2-31.8) to 42.9% (95% CI 42.5-43.2), and negative predictive value ranged from 77.6% (95% CI 77.2-77.9) to 84.9% (95% CI 84.7-85.2). Sensitivity was higher in older adults and Black patients; specificity was higher in younger adults, female patients, and White patients. Overall, sensitivity was higher for those living in areas of higher socioeconomic deprivation, while specificity was lower. These effects plateaued in the 2 least deprived quintiles of deprivation, which were comparable in both sensitivity and specificity. Predictions varied by age, with 96.2% (58,951/61,281) of those aged 70 to 79 predicted to have hypertension, whereas 0.08% of those aged 20 to 39 were predicted to have the condition. The model's performance was comparable with a more interpretable logistic regression model.

CONCLUSIONS: Despite the model's relatively good performance for those without hypertension, the positive predictive value was low, and a significant proportion of true cases remained undetected. Furthermore, there was considerable variation in performance associated with demographic characteristics, suggesting tailored approaches to case-finding may be beneficial in ensuring equity across demographic groups. Especially given the importance of understanding possible biases in predictive models, we recommend that, where there is no loss in performance, more parsimonious, transparent models be selected for prediction in health care settings. The findings of this evaluation can guide the practical application of the model, inform enhancements, direct targeted screening initiatives, and support cost-benefit analyses for broader implementation to improve hypertension management.

PMID:42743441 | DOI:10.2196/87084