Araştırma Makalesi

Evaluating Minority-Class Detection in Imbalanced Hate Speech Data via GAN-Based Synthetic Text Generation and LSHADE Hyperparameter Optimization

Cilt: 14 Sayı: 2 30 Haziran 2026
PDF İndir
EN TR

Evaluating Minority-Class Detection in Imbalanced Hate Speech Data via GAN-Based Synthetic Text Generation and LSHADE Hyperparameter Optimization

Öz

The rapid spread of hate speech on social media makes automatic detection systems indispensable; however, severe class imbalance in real-world datasets substantially limits the effectiveness of deep learning models. This study proposes an end-to-end framework that integrates a Long Short-Term Memory (LSTM) generator and a one-dimensional Convolutional Neural Network (1D-CNN) discriminator within aWGAN-GP architecture to address the imbalance problem in the Davidson hate speech dataset, where the minority class represents only 5.77% of all samples. To mitigate the training instability of standard GANs and the high computational cost of conventional grid search, the LSHADE metaheuristic is employed to autonomously optimize the main GAN hyperparameters. Experimental findings indicate that the LSHADE-optimized WGAN-GP learns the minority-class distribution without evident overfitting and produces high-quality synthetic samples that enrich the training data. When the augmented dataset is used for downstream classification, the test Macro-F1 score improves from 0.6411 to 0.6550 compared with the baseline trained on the original data only. Moreover, the number of correctly identified minority-class instances increases from 16 to 25. Multi-run evaluation with three random seeds yields very low standard deviations, especially 0.0019 on the test set, confirming the robustness and stability of the proposed pipeline.

Anahtar Kelimeler

Destekleyen Kurum

Gazi University

Etik Beyan

Ethics Approval and Consent to Participate: The study uses a publicly available secondary dataset and does not involve direct interaction with human participants. Because the work includes the synthetic generation of hate-speech-like text, generated outputs were handled only within a controlled research setting and are reported at aggregate level in order to reduce misuse risk. Availability of Data and Materials: The Davidson hate speech dataset used in this study is publicly available through the source reported in the cited literature. To support reproducibility, the manuscript explicitly reports the split protocol, preprocessing procedure, optimization objective, best hyperparameter configuration, and per-split evaluation summaries. The present submission does not include a full software release, but the experimental workflow is described in a configuration-driven manner to facilitate independent replication. Competing Interests: The authors declare no competing interests. Funding: This research received no external funding. Authors’ Contributions: Bekir Demirağ contributed to conceptualization, methodology, software development, experiments, and original draft preparation. Yusuf Sönmez and Hamdi Tolga Kahraman contributed to supervision, review, and manuscript editing.

Teşekkür

The authors thank the Department of Computer Engineering at Gazi University for its academic support during the preparation of this study.

Kaynakça

  1. [1] T. T. Nguyen, X. Yue, H. Mane, K. Seelman, P. S. P. Mullaputi, E. Dennard, J. S. Merchant, S. Criss, Y. Hswen, and Q. C. Nguyen, “Decoding digital discourse through multimodal text and image machine learning models to classify sentiment and detect hate speech in race- and LGBTQIA-related posts on social media: Quantitative study,” Journal of Medical Internet Research, vol. 27, e70053, 2025.
  2. [2] Y. Yadav, P. Bajaj, R. K. Gupta, and R. Sinha, “A comparative study of deep learning methods for hate speech and offensive language detection in textual data,” in 2021 IEEE 18th IndiaCouncil International Conference (INDICON), 2021, pp. 1–6.
  3. [3] O. Baydogan and B. Alatas, “Metaheuristic ant lion and moth flame optimization-based novel approach for automatic detection of hate speech in online social networks,” IEEE Access, vol. 9, 2021, pp. 110047–110062, doi: 10.1109/ACCESS.2021.3102596.
  4. [4] A. Luque, A. Carrasco, A. Martin, and A. de las Heras, “The impact of class imbalance in classification performance metrics based on the binary confusion matrix,” Pattern Recognition, vol. 91, 2019, pp. 216–231, doi: 10.1016/j.patcog.2019.02.023.
  5. [5] N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer, “SMOTE: Synthetic minority over-sampling technique,” Journal of Artificial Intelligence Research, vol. 16, 2002, pp. 321–357, doi: 10.1613/jair.953.
  6. [6] H. He, Y. Bai, E. A. Garcia, and S. Li, “ADASYN: Adaptive synthetic sampling approach for imbalanced learning,” in 2008 IEEE International Joint Conference on Neural Networks, 2008, pp. 1322–1328, doi: 10.1109/IJCNN.2008.4633969.
  7. [7] Ding, H., Sun, Y., Wang, Z., Huang, N., Shen, Z., and Cui, X., “RGAN-EL: A GAN and ensemble learning-based hybrid approach for imbalanced data classification,” Information Processing & Management, Vol. 60, 2023, 103235, doi: 10. 1016/j.ipm.2022.103235.
  8. [8] Ahsan, M. M., Ali, M. S., and Siddique, Z., “En-hancing and improving the performance of im¬balanced class data using novel GBO and SSG: A comparative analysis,” Neural Networks, Vol. 173, 2024, 106157, doi: 10.1016/j.neunet.2024.106157.

Ayrıntılar

Birincil Dil

İngilizce

Konular

Bilgi Sistemleri (Diğer)

Bölüm

Araştırma Makalesi

Erken Görünüm Tarihi

10 Haziran 2026

Yayımlanma Tarihi

30 Haziran 2026

Gönderilme Tarihi

2 Mayıs 2026

Kabul Tarihi

1 Haziran 2026

Yayımlandığı Sayı

Yıl 2026 Cilt: 14 Sayı: 2

Kaynak Göster

APA
Demirağ, B., Sönmez, Y., & Kahraman, H. T. (2026). Evaluating Minority-Class Detection in Imbalanced Hate Speech Data via GAN-Based Synthetic Text Generation and LSHADE Hyperparameter Optimization. Gazi Üniversitesi Fen Bilimleri Dergisi Part C: Tasarım ve Teknoloji, 14(2), 760-773. https://doi.org/10.29109/gujsc.1942853
AMA
1.Demirağ B, Sönmez Y, Kahraman HT. Evaluating Minority-Class Detection in Imbalanced Hate Speech Data via GAN-Based Synthetic Text Generation and LSHADE Hyperparameter Optimization. GUJS Part C. 2026;14(2):760-773. doi:10.29109/gujsc.1942853
Chicago
Demirağ, Bekir, Yusuf Sönmez, ve Hamdi Tolga Kahraman. 2026. “Evaluating Minority-Class Detection in Imbalanced Hate Speech Data via GAN-Based Synthetic Text Generation and LSHADE Hyperparameter Optimization”. Gazi Üniversitesi Fen Bilimleri Dergisi Part C: Tasarım ve Teknoloji 14 (2): 760-73. https://doi.org/10.29109/gujsc.1942853.
EndNote
Demirağ B, Sönmez Y, Kahraman HT (01 Haziran 2026) Evaluating Minority-Class Detection in Imbalanced Hate Speech Data via GAN-Based Synthetic Text Generation and LSHADE Hyperparameter Optimization. Gazi Üniversitesi Fen Bilimleri Dergisi Part C: Tasarım ve Teknoloji 14 2 760–773.
IEEE
[1]B. Demirağ, Y. Sönmez, ve H. T. Kahraman, “Evaluating Minority-Class Detection in Imbalanced Hate Speech Data via GAN-Based Synthetic Text Generation and LSHADE Hyperparameter Optimization”, GUJS Part C, c. 14, sy 2, ss. 760–773, Haz. 2026, doi: 10.29109/gujsc.1942853.
ISNAD
Demirağ, Bekir - Sönmez, Yusuf - Kahraman, Hamdi Tolga. “Evaluating Minority-Class Detection in Imbalanced Hate Speech Data via GAN-Based Synthetic Text Generation and LSHADE Hyperparameter Optimization”. Gazi Üniversitesi Fen Bilimleri Dergisi Part C: Tasarım ve Teknoloji 14/2 (01 Haziran 2026): 760-773. https://doi.org/10.29109/gujsc.1942853.
JAMA
1.Demirağ B, Sönmez Y, Kahraman HT. Evaluating Minority-Class Detection in Imbalanced Hate Speech Data via GAN-Based Synthetic Text Generation and LSHADE Hyperparameter Optimization. GUJS Part C. 2026;14:760–773.
MLA
Demirağ, Bekir, vd. “Evaluating Minority-Class Detection in Imbalanced Hate Speech Data via GAN-Based Synthetic Text Generation and LSHADE Hyperparameter Optimization”. Gazi Üniversitesi Fen Bilimleri Dergisi Part C: Tasarım ve Teknoloji, c. 14, sy 2, Haziran 2026, ss. 760-73, doi:10.29109/gujsc.1942853.
Vancouver
1.Bekir Demirağ, Yusuf Sönmez, Hamdi Tolga Kahraman. Evaluating Minority-Class Detection in Imbalanced Hate Speech Data via GAN-Based Synthetic Text Generation and LSHADE Hyperparameter Optimization. GUJS Part C. 01 Haziran 2026;14(2):760-73. doi:10.29109/gujsc.1942853

                                     16168      16167     16166     21432        logo.png   


    e-ISSN:2147-9526