J Vis Exp. 2026 Jul 7;(233). doi: 10.3791/69857.
ABSTRACT
Heart disease remains one of the leading causes of mortality worldwide, creating an urgent need for accurate and scalable predictive systems that enable early diagnosis and timely clinical intervention. Traditional machine learning approaches often struggle to efficiently process large-scale medical datasets and lack interpretability, limiting their usefulness for supporting clinical decision-making. To address these challenges, this study proposes a Cluster Visualized Distributed Machine Learning framework for heart disease prediction. The framework incorporates two distributed algorithms: Cluster Visualized Hadoop Distributed Decision Tree (CViHDDT) and Cluster Visualized Hadoop Distributed K-Nearest Neighbor (CViHDKNN). The proposed models leverage Hadoop's MapReduce framework for distributed computation across large datasets, while integrating K-Means clustering for improved data organization and visualization. This cluster-based visualization enhances interpretability by allowing clinicians to better understand relationships among patient risk factors and prediction outcomes. Experimental evaluation was conducted using the UCI Heart Disease dataset in a Hadoop-based distributed environment. The results show that CViHDKNN achieved superior predictive performance, achieving 85.25% accuracy and 88% recall, outperforming the CViHDDT model, which achieved 80.33% accuracy. Adjusting classification cut-off values also influenced sensitivity and detection rates: lower cut-offs improved true-positive detection while maintaining acceptable false-positive levels. These findings demonstrate that clustering-enhanced distributed learning improves scalability, predictive accuracy, and clinical interpretability for heart disease prediction.
PMID:42507720 | DOI:10.3791/69857