Abstract
Introduction: Metabolic syndrome (MetS) is a complex cluster of interconnected metabolic irregularities that seriously increase the risk of cardiovascular disorders and type 2 diabetes mellitus. Forecasting MetS and its associated factors with machine-learning (ML) models offers a promising approach to analyze datasets and uncover patterns associated with MetS risk. Accordingly, this study examined the performance of ML models in predicting MetS using distinctive combinations of feature selection and normalization techniques.
Methods: To this end, decision tree, K-nearest neighbors (KNN), Naïve Bayes, neural networks, random forest (RF), and support vector machine (SVM) were employed, evaluating their accuracy, precision, sensitivity, and specificity. Moreover, feature selection methods included the Chi2 score, F score, and mutual information, while normalization techniques comprised MinMax, StandardScaler (STD), and Robust.
Results: Our analysis revealed that ensemble models, particularly those combining SVM, RF, and KNN with mutual information and STD normalization, outperformed individual models in terms of accuracy and F1 scores, reaching an accuracy of 0.85% and an impressive F1 score of 0.93. Furthermore, the results revealed common features, such as age, body mass index, waist circumference, high blood pressure, diabetes, and specific nutritional intakes per day that consistently influenced MetS risk across different models.
Conclusion: Overall, our findings highlight the efficacy of ML models in predicting MetS and the importance of considering both clinical and lifestyle factors in predictive modeling. Using large datasets and advanced algorithms, healthcare practitioners can identify individuals at high risk of MetS and implement interventions to moderate their risk and improve health outcomes.