Research Article

Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification

Volume: 30 Number: 2 August 25, 2026
TR EN

Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification

Abstract

Abstract: Class imbalance constitutes a significant challenge in text classification tasks, particularly due to the insufficient representation of minority-class instances. Traditional machine learning methods trained on imbalanced datasets tend to exhibit a bias toward majority classes, thereby limiting their ability to accurately classify minority-class samples. To address this issue, this study proposes a classification framework, termed Imb-LLMClass, to investigate the effectiveness of large language model-based classification approaches for imbalanced text data. Experimental results demonstrate that the proposed Imb-LLMClass framework achieved the highest performance among the compared methods when evaluated based on the average of three experimental runs, attaining a Macro F1-Score of 0.878 and a Balanced Accuracy of 0.870 across three datasets. Furthermore, the increase in the average minority-class recall to 0.834 indicates that the proposed approach not only improves overall classification performance but also enhances the learning of underrepresented classes. These findings indicate that, in imbalanced text classification tasks, evaluating large language models solely based on overall performance metrics may be insufficient and that their ability to effectively represent minority classes should also be taken into consideration.

Keywords

References

  1. [1] Susan, S., Kumar, A. 2021. The balancing trick: Optimized sampling of imbalanced datasets—A brief survey of the recent state of the art. Engineering Reports, 3(4).
  2. [2] Aubaidan, B. H., Kadir, R. A., Lajb, M. T., Anwar, M., Qureshi, K. N., Taha, B. A., Ghafoor, K. 2025. A review of intelligent data analysis: Machine learning approaches for addressing class imbalance in healthcare - challenges and perspectives. Intelligent Data Analysis: An International Journal, 29(3), 699-719.
  3. [3] Chawla, N. V., Bowyer, K. W., Hall, L. O., Kegelmeyer, W. P. 2002. SMOTE: Synthetic Minority Over-sampling Technique. Journal of Artificial Intelligence Research, 16, 321-357.
  4. [4] Li, X., Roth, D. 2002. Learning question classifiers. Proceedings of the 19th International Conference on Computational Linguistics, 1-7.
  5. [5] Saravia, E., Liu, H. C. T., Huang, Y. H., Wu, J., Chen, Y. S. 2018. CARER: Contextualized Affect Representations for Emotion Recognition. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 3687-3697.
  6. [6] Çöltekin, Ç. 2020. A corpus of Turkish offensive language on social media. Proceedings of the 12th Language Resources and Evaluation Conference (LREC 2020), 6174-6184, Marseille, France. European Language Resources Association. URL: https://aclanthology.org/2020.lrec-1.758/
  7. [7] Indrawati, A., Subagyo, H., Sihombing, A., Wagiyah, W., Afandi, S. 2020. Analyzing the impact of resampling method for imbalanced data text in Indonesian scientific articles categorization. BACA: Jurnal Dokumentasi dan Informasi, 41(2), 133-141.
  8. [8] Mahmoudi, L., Salem, M., Alharbe, N. R. 2026. Addressing class imbalance in text classification with LLMs: A prompt-based GPT-2 approach. Journal of Information Science , 52(3), 908-922.

Details

Primary Language

English

Subjects

Artificial Intelligence (Other)

Journal Section

Research Article

Publication Date

August 25, 2026

Submission Date

June 3, 2026

Acceptance Date

August 12, 2026

Published in Issue

Year 2026 Volume: 30 Number: 2

APA
Tan, F. G. (2026). Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification. Süleyman Demirel Üniversitesi Fen Bilimleri Enstitüsü Dergisi, 30(2), 224-238. https://doi.org/10.19113/sdufenbed.1963184
AMA
1.Tan FG. Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification. J. Nat. Appl. Sci. 2026;30(2):224-238. doi:10.19113/sdufenbed.1963184
Chicago
Tan, Fatma Gülşah. 2026. “Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification”. Süleyman Demirel Üniversitesi Fen Bilimleri Enstitüsü Dergisi 30 (2): 224-38. https://doi.org/10.19113/sdufenbed.1963184.
EndNote
Tan FG (August 1, 2026) Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification. Süleyman Demirel Üniversitesi Fen Bilimleri Enstitüsü Dergisi 30 2 224–238.
IEEE
[1]F. G. Tan, “Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification”, J. Nat. Appl. Sci., vol. 30, no. 2, pp. 224–238, Aug. 2026, doi: 10.19113/sdufenbed.1963184.
ISNAD
Tan, Fatma Gülşah. “Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification”. Süleyman Demirel Üniversitesi Fen Bilimleri Enstitüsü Dergisi 30/2 (August 1, 2026): 224-238. https://doi.org/10.19113/sdufenbed.1963184.
JAMA
1.Tan FG. Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification. J. Nat. Appl. Sci. 2026;30:224–238.
MLA
Tan, Fatma Gülşah. “Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification”. Süleyman Demirel Üniversitesi Fen Bilimleri Enstitüsü Dergisi, vol. 30, no. 2, Aug. 2026, pp. 224-38, doi:10.19113/sdufenbed.1963184.
Vancouver
1.Fatma Gülşah Tan. Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification. J. Nat. Appl. Sci. 2026 Aug. 1;30(2):224-38. doi:10.19113/sdufenbed.1963184

e-ISSN :1308-6529
Linking ISSN (ISSN-L): 1300-7688

All published articles in the journal can be accessed free of charge and are open access under the Creative Commons CC BY-NC (Attribution-NonCommercial) license. All authors and other journal users are deemed to have accepted this situation. Click here to access detailed information about the CC BY-NC license.