A Comparative Performance Evaluation of Classification Algorithms on Imbalanced Datasets
Abstract
Class imbalance remains a critical challenge in supervised learning, often biasing classifiers toward majority classes. While resampling techniques like Synthetic Minority Oversampling Technique (SMOTE) are widely used, the combined effect of data balancing and hyperparameter optimization across diverse datasets is rarely systematically explored. This study presents a comprehensive comparative analysis of four classification algorithms—Naive Bayes (NB), K-Nearest Neighbors (K-NN), Artificial Neural Networks (ANN), and Random Forest (RF)—across ten benchmark datasets from the UCI Machine Learning Repository. Unlike previous studies relying on default parameters, this research employs a rigorous Grid Search strategy to optimize hyperparameters for each algorithm within a rigorous SMOTE-balanced stratified cross-validation pipeline to ensure robust evaluation. Performance was assessed using a wide range of metrics, including Accuracy, Precision, Recall, F1-score, and Area Under the Curve (AUC). Experimental results reveal that ANN achieved the highest robustness in high-dimensional and complex categorical datasets (e.g., Car Evaluation F1-score: 0.990), significantly outperforming traditional models. Conversely, RF demonstrated superior stability in datasets with high feature dimensionality (e.g., Arrhythmia F1-score: 0.600) and chemical interactions (e.g., QSAR Fish Toxicity F1-score: 0.839). While K-NN remained competitive in low-dimensional spaces, NB struggled with complex feature dependencies. This study contributes to the literature by demonstrating that algorithmic superiority is context-dependent and providing a data-driven framework for selecting classifiers based on structural characteristics such as dimensionality, categorical complexity, and sample size.
Keywords
- Imbalanced data
- Classification algorithms
- Naive Bayes
- K-Nearest Neighbors
- Artificial Neural Networks
- SMOTE
- Performance evaluation
Ethical Statement
Thanks
References
- A. Tharwat, “Classification assessment methods,” Appl. Comput. Informat., vol. 17, no. 1, pp. 168–192, 2021, doi: 10.1016/j.aci.2018.08.003. [Online]. Available: https://doi.org/10.1016/j.aci.2018.08.003
- T. Saito and M. Rehmsmeier, “The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets,” PLOS ONE, vol. 10, no. 3, Art. no. e0118432, 2015, doi: 10.1371/journal.pone.0118432. [Online]. Available: https://doi.org/10.1371/journal.pone.0118432
- N. V. Chawla, N. Japkowicz, and A. Kotcz, “Editorial: Special issue on learning from imbalanced data sets,” ACM SIGKDD Explor. Newslett., vol. 6, no. 1, pp. 1–6, 2004, doi: 10.1145/1007730.1007733. [Online]. Available: https://doi.org/10.1145/1007730.1007733
- J. M. Johnson and T. M. Khoshgoftaar, “Survey on deep learning with class imbalance,” J. Big Data, vol. 6, no. 1, Art. no. 27, 2019, doi: 10.1186/s40537-019-0192-5. [Online]. Available: https://doi.org/10.1186/s40537-019-0192-5
- W. Chen, K. Yang, Z. Yu, Y. Shi, and C. L. P. Chen, “A survey on imbalanced learning: Latest research, applications and future directions,” Artif. Intell. Rev., vol. 57, Art. no. 137, 2024, doi: 10.1007/s10462-024-10759-6. [Online]. Available: https://doi.org/10.1007/s10462-024-10759-6
- M. Salmi et al., “Handling imbalanced medical datasets: Review of a decade of research,” Artif. Intell. Rev., vol. 57, Art. no. 273, 2024, doi: 10.1007/s10462-024-10884-2. [Online]. Available: https://doi.org/10.1007/s10462-024-10884-2
- H. Kaur, H. S. Pannu, and A. K. Malhi, “A systematic review on imbalanced data challenges in machine learning: Applications and solutions,” ACM Comput. Surv., vol. 52, no. 4, pp. 1–36, 2019, doi: 10.1145/3343440. [Online]. Available: https://doi.org/10.1145/3343440
- N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “SMOTE: Synthetic minority over-sampling technique,” J. Artif. Intell. Res., vol. 16, pp. 321–357, 2002, doi: 10.1613/jair.953. [Online]. Available: https://doi.org/10.1613/jair.953
Details
Primary Language
English
Subjects
Artificial Intelligence (Other)
Journal Section
Research Article
Publication Date
September 30, 2026
Submission Date
December 1, 2025
Acceptance Date
February 3, 2026
Published in Issue
Year 2026 Volume: 9 Number: 4