Araştırma Makalesi

Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification

Cilt: 30 Sayı: 2 25 Ağustos 2026
PDF İndir
TR EN

Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification

Öz

Abstract: Class imbalance constitutes a significant challenge in text classification tasks, particularly due to the insufficient representation of minority-class instances. Traditional machine learning methods trained on imbalanced datasets tend to exhibit a bias toward majority classes, thereby limiting their ability to accurately classify minority-class samples. To address this issue, this study proposes a classification framework, termed Imb-LLMClass, to investigate the effectiveness of large language model-based classification approaches for imbalanced text data. Experimental results demonstrate that the proposed Imb-LLMClass framework achieved the highest performance among the compared methods when evaluated based on the average of three experimental runs, attaining a Macro F1-Score of 0.878 and a Balanced Accuracy of 0.870 across three datasets. Furthermore, the increase in the average minority-class recall to 0.834 indicates that the proposed approach not only improves overall classification performance but also enhances the learning of underrepresented classes. These findings indicate that, in imbalanced text classification tasks, evaluating large language models solely based on overall performance metrics may be insufficient and that their ability to effectively represent minority classes should also be taken into consideration.

Anahtar Kelimeler

Kaynakça

  1. [1] Susan, S., Kumar, A. 2021. The balancing trick: Optimized sampling of imbalanced datasets—A brief survey of the recent state of the art. Engineering Reports, 3(4).
  2. [2] Aubaidan, B. H., Kadir, R. A., Lajb, M. T., Anwar, M., Qureshi, K. N., Taha, B. A., Ghafoor, K. 2025. A review of intelligent data analysis: Machine learning approaches for addressing class imbalance in healthcare - challenges and perspectives. Intelligent Data Analysis: An International Journal, 29(3), 699-719.
  3. [3] Chawla, N. V., Bowyer, K. W., Hall, L. O., Kegelmeyer, W. P. 2002. SMOTE: Synthetic Minority Over-sampling Technique. Journal of Artificial Intelligence Research, 16, 321-357.
  4. [4] Li, X., Roth, D. 2002. Learning question classifiers. Proceedings of the 19th International Conference on Computational Linguistics, 1-7.
  5. [5] Saravia, E., Liu, H. C. T., Huang, Y. H., Wu, J., Chen, Y. S. 2018. CARER: Contextualized Affect Representations for Emotion Recognition. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, 3687-3697.
  6. [6] Çöltekin, Ç. 2020. A corpus of Turkish offensive language on social media. Proceedings of the 12th Language Resources and Evaluation Conference (LREC 2020), 6174-6184, Marseille, France. European Language Resources Association. URL: https://aclanthology.org/2020.lrec-1.758/
  7. [7] Indrawati, A., Subagyo, H., Sihombing, A., Wagiyah, W., Afandi, S. 2020. Analyzing the impact of resampling method for imbalanced data text in Indonesian scientific articles categorization. BACA: Jurnal Dokumentasi dan Informasi, 41(2), 133-141.
  8. [8] Mahmoudi, L., Salem, M., Alharbe, N. R. 2026. Addressing class imbalance in text classification with LLMs: A prompt-based GPT-2 approach. Journal of Information Science , 52(3), 908-922.

Ayrıntılar

Birincil Dil

İngilizce

Konular

Yapay Zeka (Diğer)

Bölüm

Araştırma Makalesi

Yayımlanma Tarihi

25 Ağustos 2026

Gönderilme Tarihi

3 Haziran 2026

Kabul Tarihi

12 Ağustos 2026

Yayımlandığı Sayı

Yıl 2026 Cilt: 30 Sayı: 2

Kaynak Göster

APA
Tan, F. G. (2026). Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification. Süleyman Demirel Üniversitesi Fen Bilimleri Enstitüsü Dergisi, 30(2), 224-238. https://doi.org/10.19113/sdufenbed.1963184
AMA
1.Tan FG. Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification. Süleyman Demirel Üniv. Fen Bilim. Enst. Derg. 2026;30(2):224-238. doi:10.19113/sdufenbed.1963184
Chicago
Tan, Fatma Gülşah. 2026. “Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification”. Süleyman Demirel Üniversitesi Fen Bilimleri Enstitüsü Dergisi 30 (2): 224-38. https://doi.org/10.19113/sdufenbed.1963184.
EndNote
Tan FG (01 Ağustos 2026) Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification. Süleyman Demirel Üniversitesi Fen Bilimleri Enstitüsü Dergisi 30 2 224–238.
IEEE
[1]F. G. Tan, “Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification”, Süleyman Demirel Üniv. Fen Bilim. Enst. Derg., c. 30, sy 2, ss. 224–238, Ağu. 2026, doi: 10.19113/sdufenbed.1963184.
ISNAD
Tan, Fatma Gülşah. “Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification”. Süleyman Demirel Üniversitesi Fen Bilimleri Enstitüsü Dergisi 30/2 (01 Ağustos 2026): 224-238. https://doi.org/10.19113/sdufenbed.1963184.
JAMA
1.Tan FG. Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification. Süleyman Demirel Üniv. Fen Bilim. Enst. Derg. 2026;30:224–238.
MLA
Tan, Fatma Gülşah. “Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification”. Süleyman Demirel Üniversitesi Fen Bilimleri Enstitüsü Dergisi, c. 30, sy 2, Ağustos 2026, ss. 224-38, doi:10.19113/sdufenbed.1963184.
Vancouver
1.Fatma Gülşah Tan. Imb-LLMClass: A Large Language Model-Based Framework for Imbalanced Text Classification. Süleyman Demirel Üniv. Fen Bilim. Enst. Derg. 01 Ağustos 2026;30(2):224-38. doi:10.19113/sdufenbed.1963184

e-ISSN :1308-6529
Linking ISSN (ISSN-L): 1300-7688

Dergide yayımlanan tüm makalelere ücretiz olarak erişilebilinir ve Creative Commons CC BY-NC Atıf-GayriTicari lisansı ile açık erişime sunulur. Tüm yazarlar ve diğer dergi kullanıcıları bu durumu kabul etmiş sayılırlar. CC BY-NC lisansı hakkında detaylı bilgiye erişmek için tıklayınız.