Authors: Dr Chunnu Lal, Dr Satender Kumar, Dr Raj kumar
Abstract: The prevalence of Diabetes Mellitus, a chronic metabolic disorder, is rapidly increasing worldwide, which is creating an urgent need for early risk assessment tools that are fast, accessible, and accurate. This research presents a comprehensive study of a diabetes prediction web application that combines three heterogeneous machine learning classifiers—Random Forest, Support Vector Machine (SVM), and Logistic Regression—in a soft-voting ensemble architecture. The Pima Indians Diabetes Database (PIDD) was used to train and validate the integrated system, which uses eight clinically relevant health parameters for binary classification. The ensemble model's test set accuracy was 81.04%, which is a significant improvement over the individual base classifiers (LR: 76.60%, SVM: 75.32%, RF: 78.45%). The system demonstrated strong clinical utility with a precision of 77.6% and a recall rate of 65.0%, establishing a critical balance for medical screening applications. By deploying the complete architecture through a lightweight Flask-based web interface, it was possible to predict risks in real-time with millisecond-level inference latency. This work validates the effectiveness of soft-voting ensemble methodologies in achieving robust classification accuracy for high-stakes healthcare applications and demonstrates a scalable, practical implementation pattern for deploying machine learning models to address significant public health challenges.
International Journal of Science, Engineering and Technology