Authors: Lalit Verma, Manjot Kaur, Shivani Sharma, Siya Thakur
Abstract: Cardiovascular disease is one of the leading causes of mortality worldwide, which makes early and accurate prediction essential for timely clinical intervention. This study conducted a comparative analysis of six machine learning models, namely Random Forest, Gradient Boosting, Naive Bayes, Support Vector Machine (SVM), K-Nearest Neighbors (KNN), and Light Gradient Boosting Machine (LightGBM), for the prediction of heart disease from clinical data. Each model was trained and evaluated on a clinical dataset comprising established cardiovascular risk factors, and performance was assessed using accuracy, precision, recall, F1-score, and confusion matrix analysis. Naive Bayes achieved the highest accuracy at 90.74 percent, followed by SVM at 88.88 percent and LightGBM at 83.33 percent. Random Forest, KNN, and Gradient Boosting achieved accuracies of 79.62 percent, 81.48 percent, and 77.77 percent respectively. The results show that probabilistic classifiers performed particularly well on the relatively small clinical dataset used in this study, whereas ensemble tree based methods required larger training samples to fully exploit their variance reduction capability. Naive Bayes also produced the best recall and F1-score, while SVM achieved the highest precision. These findings indicate that model selection for heart disease prediction should be guided by dataset characteristics and by the clinical priorities of the intended application rather than by model complexity alone. The comparative framework developed in this study provides a practical basis for selecting machine learning models suited to cardiovascular decision support systems and highlights directions for future work involving larger datasets and deep learning approaches.
International Journal of Science, Engineering and Technology