skip to main content

KLASIFIKASI MENGGUNAKAN ALGORITMA K-NEAREST NEIGHBOR DAN C5.0 PADA IMBALANCE CLASS DATA DENGAN SMOTE

*Salsabilla Rizka Ardhana  -  Departemen Statistika, Fakultas Sains dan Matematika, Universitas Diponegoro, Indonesia
Tatik Widiharih  -  Departemen Statistika, Fakultas Sains dan Matematika, Universitas Diponegoro, Indonesia
Bagus Arya Saputra  -  Departemen Statistika, Fakultas Sains dan Matematika, Universitas Diponegoro, Indonesia
Open Access Copyright 2026 Jurnal Gaussian under http://creativecommons.org/licenses/by-nc-sa/4.0.

Citation Format:
Abstract

Rural Banks (BPR) provide financial services to micro-businesses and low repayment communities, especially in rural areas. The main activity of the bank is lending. Customer credit classification is expected to assist BPR in anticipating potentially bad loans. K-Nearest Neighbor and C5.0 classify current and potentially bad credit status based on customer data from BPR “X” in Central Java in October 2022. K-Nearest Neighbor is effective against a large amount of training data and works based on the nearest neighbor. C5.0 can improve classification accuracy and work by calculating entropy, gain, split info, and gain ratio to form a decision tree. There is an imbalance class data which causes the classification process to focus more on the majority class. Imbalance class data is handled using SMOTE as an oversampling approach. Classification with the addition of SMOTE can improve the evaluation of classification accuracy, especially G-Mean. G-mean is the most comprehensive measurement compared to accuracy, sensitivity and specificity in evaluating classification performance on imbalance class data. Result show an increased G-Mean to 58.55% on KNN and 64.05% on C5.0. Based on the classification results, it is concluded that C5.0 with SMOTE is a more appropriate classification model for customer credit status.

Keywords: Credit Status; K-Nearest Neighbor; C5.0; Imbalance Class Data; SMOTE

Article Metrics:

Article Info
Section: Articles
Language : EN
  1. Chawla, N. V., Bowyer, K. W., Hall, L. O., dan Kegelmeyer, W.P. 2002. SMOTE: Synthetic Minority Over-Sampling Technique. Journal of Artificial Intelligence Research, 16, 321-357
  2. Han, J., dan Kamber, M. 2006. Data Mining Concepts and Techniques Second Edition. San Fransisco: Morgan Kaufmann
  3. Hasan, N. I. (2014). Pengantar Perbankan. Jakarta: Referensi (Gaung Persada Press Group)
  4. He, H., dan Gracia, E.A. 2009. Learning from Imbalanced Data, IEEE Trans. Knowl. Discov. 21(9) 1263-1284
  5. Kantardzic, M. 2011. Data Mining: Concept, Models, Methods, and Algorithms. 2nd edition. New Jersey: John W & Sons, Inc
  6. Kubat, M., Holte, R., dan Matwin, S. 1997. Learning When Negative Examples Abound. In European conference on machine learning (pp. 146-153). Springer, Berlin, Heidelberg
  7. Prasetyo, E. 2012. Data Mining Konsep dan Aplikasi Menggunakan MATLAB. Yogyakarta: ANDI Yogyakarta
  8. Singh, P. dan Sharma, P. A. 2019. Analysis of Imbalanced Classification Algorithms: A Perspective View. International Journal of Trend in Scientific Research and Development, 3(2), 974-978
  9. Sreemathy, J., dan Balamurugan, P. S. (2012). An Efficient Text Classification using KNN and Naïve Bayesian. International Journal on Computer Science and Engineering, 4(3), 392
  10. Tan, P., Steinbach, M., dan Kumar, V. 2006. Introduction to Data Mining. Boston: Pearson Education

Last update:

No citation recorded.

Last update:

No citation recorded.