COMPARATIVE ANALYSIS OF RANDOM FOREST AND SUPPORT VECTOR MACHINE ALGORITHMS FOR DIABETES MELLITUS PREDICTION
Abstract
Abstract. Early prediction of type 2 diabetes mellitus is important to facilitate faster and more accurate diagnosis and clinical decision-making. This study aims to compare the performance of the Random Forest and Support Vector Machine (SVM) algorithms in predicting diabetes and to analyze the effect of applying the Synthetic Minority Over-sampling Technique (SMOTE) to imbalanced data. The study used the Bangladesh Diabetes 2025 dataset, following the stages of data selection, preprocessing, data transformation, modeling, and evaluation. The preprocessing stage included median imputation and the removal of duplicate data, while feature standardization was performed after data splitting to prevent data leakage. The dataset was split using an 80:20 ratio with a stratification technique. The study applied two experimental scenarios: one without SMOTE and one with SMOTE applied only to the training data. Model evaluation was conducted using the metrics accuracy, precision, recall, and F1-score based on the test data. The results show that Random Forest outperforms SVM. In the scenario without SMOTE, Random Forest achieved an accuracy of 94.8%, precision of 96.4%, recall of 97.0%, and an F1-score of 96.7%, while SVM achieved an accuracy of 91.5% and an F1-score of 94.6%. After applying SMOTE, Random Forest’s performance improved slightly to an accuracy of 95.3% and an F1-score of 97.0%, while SVM’s performance declined to an accuracy of 90.6% and an F1-score of 93.9%. The results of the study show that Random Forest is the best model for predicting type 2 diabetes mellitus in the dataset used.
Downloads
References
International Diabetes Federation, “IDF Diabetes Atlas (11th ed.),” 2025. Accessed: Dec. 07, 2025. [Online]. Available: https://diabetesatlas.org/
A. Renata Agusti, A. Fauzi, K. Ahmad Baihaqi, and T. Rohana, “Identifikasi Penyakit Diabetes Mellitus Menggunakan Algoritma Support Vector Machine dan Random Forest,” Jurnal Riset Komputer), vol. 12, no. 4, pp. 2407–389, 2025, doi: 10.30865/jurikom.v12i4.8686.
M. Fadli Kurniawan and D. Ayu Megawaty, “Comparison of Logistic Regression, Random Forest, Support Vector Machine (SVM) and K-Nearest Neighbor (KNN) Algorithms in Diabetes Prediction,” 2025. [Online]. Available: http://jurnal.polibatam.ac.id/index.php/JAIC
S. I. Fallo and A. Nay, “KLASIFIKASI SUPPORT VECTOR MACHINE DAN RANDOM FOREST PADA DATA BIOMEDIS : APLIKASI DALAM ANALISIS DATA PENYAKIT DIABETES CLASSIFICATION OF SUPPORT VECTOR MACHINE AND RANDOM FOREST ON BIOMEDICAL DATA: APPLICATIONS IN DIABETES DISEASE DATA ANALYSIS,” 2024.
H. Alrasyid, A. Homaidi, M. Kom, and Z. Fatah, “Comparison Support Vector Machine and Random Forest Algorithms in Detect Diabetes,” 2024.
M. Y. Bhuiyan, S. S. Ayon, Md. E. Hossain, M. S. U. Miah, and A. Sarower, “Type-2 Diabetes Dataset Bangladesh,” Nov. 12, 2025, Mendeley Data.
N. Novitasari, N. D. Nuris, and R. Herdiana, “Penerapan Algoritma K-Means untuk Clustering Data Jumlah Penduduk Miskin Berdasarkan Kota/Kabupaten di Jawabarat menggunakan Rapidminer,” Jurnal Informatika Terpadu, vol. 9, no. 1, pp. 68–73, Mar. 2023, doi: 10.54914/jit.v9i1.660.
M. Desiawan and A. Solichin, “SVM Optimization with Grid Search Cross Validation for Improving Accuracy of Schizophrenia Classification Based on EEG Signal,” JURNAL TEKNIK INFORMATIKA, vol. 17, no. 1, pp. 10–20, May 2024, doi: 10.15408/jti.v17i1.37422.
F. S. Pratiwi, M. Agung Barata, and A. D. Ardianti, “IMPLEMENTASI METODE SMOTE DAN RANDOM OVER-SAMPLING PADA ALGORITMA MACHINE LEARNING UNTUK PREDIKSI CUSTOMER CHURN DI SEKTOR PERBANKAN,” Jurnal Sistem Informasi dan Informatika (Simika), vol. 8, no. 1, 2025, [Online]. Available: https://www.kaggle.com/datasets/gauravtopre/bank-customer-churn-dataset/data
A. R. MOHAMMED, “Enhancing Diabetes Mellitus Onset Prediction through Advanced Ensemble Learning Techniques,” Journal of Statistical Modelling and Analytics, vol. 6, no. 2, pp. 1–18, Dec. 2024, doi: 10.22452/josma.vol6no2.2.
M. Talebi Moghaddam et al., “Predicting diabetes in adults: identifying important features in unbalanced data over a 5-year cohort study using machine learning algorithm,” BMC Med. Res. Methodol., vol. 24, no. 1, p. 220, Sep. 2024, doi: 10.1186/s12874-024-02341-z.
Y. Xu and Y. Nie, “Diabetes Prediction Based on Support Vector Machine Model,” Highlights in Science, Engineering and Technology, vol. 102, pp. 311–320, Jul. 2024, doi: 10.54097/ygdasy74.
M. Imani, M. Joudaki, A. Bagheri, and H. R. Arabnia, “Why ROC-AUC Is Misleading for Highly Imbalanced Data: In-Depth Evaluation of MCC, F2-Score, H-Measure, and AUC-Based Metrics Across Diverse Classifiers,” Technologies (Basel)., vol. 14, no. 1, p. 54, Jan. 2026, doi: 10.3390/technologies14010054.
M. A. Widyananda and I. Palupi, “Implementation of the Spiral Optimization Algorithm in the Support Vector Machine (SVM) Classification Method (Case Study: Diabetes Prediction),” in 2021 International Conference Advancement in Data Science, E-learning and Information Systems (ICADEIS), IEEE, Oct. 2021, pp. 1–6. doi: 10.1109/ICADEIS52521.2021.9701953.
O. Toyin, J.-J. Timothy, S. E. Obamiyi, B. A. O, O. M. Aweh, and A. A. James, “Comparative Analysis of Supervised Machine Learning Algorithms for Diabetes Prediction,” INTERNATIONAL JOURNAL OF MATHEMATICS AND COMPUTER RESEARCH, vol. 12, no. 08, Aug. 2024, doi: 10.47191/ijmcr/v12i8.05.
B. R. Prasetyo, E. D. Wahyuni, and P. M. Kusumantara, “KOMPARASI PERFORMA MODEL BERBASIS ALGORITMA RANDOM FOREST DAN LIGHTGBM DALAM MELAKUKAN KLASIFIKASI DIABETES MELITUS GESTASIONAL,” Jurnal Informatika dan Teknik Elektro Terapan, vol. 12, no. 3, Aug. 2024, doi: 10.23960/jitet.v12i3.4817.

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.



