Analyzing the Efficacy of K-Means Clustering and Logistic Regression For Diabetes Prediction
DOI:
https://doi.org/10.70135/seejph.vi.2454Keywords:
Artificial Neural Network, Machine Learning, Support Vector Machine, knowledge discovery in databases, Online Analytical Processing..Abstract
Diabetes causes a large number of deaths each year and a large number of people living with the disease do not realize their health condition early enough. In this study, we propose a data mining based model for early diagnosis and prediction of diabetes using the UCI database. Although K-means is simple and can be used for a wide variety of data types, it is quite sensitive to initial positions of cluster centers which determine the final cluster result, which either provides a sufficient and efficiently clustered dataset for the logistic regression model, or gives a lesser amount of data as a result of incorrect clustering of the original dataset, thereby limiting the performance of the logistic regression model. Our findings offer insights into the comparative strengths and weaknesses of each method, shedding light on their potential applications in diabetes diagnosis and risk assessment. A further experiment with a new dataset showed the applicability of our model for the predication of diabetes.
Downloads
Published
How to Cite
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NoDerivatives 4.0 International License.