A Comparative Evaluation of Three-Class Diabetes Classification Using Machine Learning Algorithms
Ayeni Taiwo Michael *
College of Professional Studies, Analytics, Northeastern University, Toronto, Canada.
Odukoya Elijah Ayooluwa
Department of Statistics, Federal University of Technology and Environmental Science, Iyin-Ekiti, Nigeria.
Ilesanmi Anthony Opeyemi
Department of Statistics, Ekiti State University, Ado-Ekiti, Nigeria.
*Author to whom correspondence should be addressed.
Abstract
Background: Early detection of diabetes and prediabetes is important for reducing long-term complications.
Aim: This study comparatively evaluated four machine learning algorithms for three-class diabetes classification using routinely available clinical indicators from the National Health and Nutrition Examination Survey.
Methods: The analytical sample included 2,029 participants classified as normal (62.0%), prediabetic (26.7%), or diabetic (11.3%). Recursive Feature Elimination with 10-fold cross-validation was used to select predictors from 27 candidate variables. Six features were retained: fasting glucose, age, diabetes history, insulin level, waist circumference, and systolic blood pressure. Multinomial Logistic Regression, Decision Trees, Random Forest, and XGBoost were trained using a 70/30 stratified train-test split and evaluated using accuracy, Cohen’s Kappa, and class-specific performance metrics.
Results: Random Forest achieved the highest overall test performance, with 78.1% accuracy and a Cohen’s Kappa of 0.471. XGBoost, Multinomial Logistic Regression, and Decision Trees achieved accuracies of 72.6%, 73.0%, and 72.4%, respectively. All models showed high specificity for diabetes detection, exceeding 97%. Prediabetes classification remained difficult across algorithms, with sensitivity ranging from 38% to 42%.
Conclusion: Random Forest provided the best overall performance for three-class diabetes classification in this analytical sample. However, modest agreement, low prediabetes sensitivity, potential label leakage from fasting glucose, and the absence of external validation indicate that further evaluation is required before clinical application.
Keywords: Diabetes classification, machine learning, Random Forest, XGBoost, Multinomial Logistic Regression, Decision Trees, NHANES, prediabetes detection, three-class classification, feature selection.