- Conference Article
- 10.1109/bmeicon66226.2025.11113699
Comparative Performance Analysis of One-Class and Binary-Class Models for Diabetes Risk Prediction
- Jul 15, 2025
- Pathrapim Panthong + 2 more +2
Diabetes poses a significant global health challenge due to its rising prevalence and the high number of undiagnosed cases. Early detection is crucial, and machine learning presents promising solutions for enhancing diagnostic accuracy. This study compares binary-class (supervised learning) and one-class classification (unsupervised learning) approaches for diabetes prediction using clinical data from the publicly available Kaggle Diabetes Dataset, which includes eight features: age, gender, hypertension, heart disease, BMI, HbA1c level, blood glucose level, and smoking history. The data is strongly biased towards diabetic cases, with 91500 cases compared to 8500 non-diabetic cases. Data preprocessing, feature selection, and model training were conducted in Python using Jupyter Notebook. For binary classification, several supervised machine learning algorithms are employed, including logistic regression, decision trees, random forests, gradient boosting, adaptive boosting, support vector machines, K-nearest neighbors, and Naïve Bayes. These models are trained and evaluated using accuracy, precision, recall, and F1-score, with 5-fold cross-validation to ensure model robustness. For one-class classification, 91,500 diabetic cases were used to train an unsupervised machine learning support vector machine model. The trained one-class model was then evaluated using both diabetic and non-diabetic cases, and it was found that the one-class model offers better and more consistent performance compared to the binary class, with overall performance metrics above 95%. The results highlight key performance differences between classification strategies, offering insights into the optimal modeling approach for clinical screening and biomedical applications.
Read more