Araştırma Makalesi

Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts

Cilt: 9 Sayı: 5 15 Eylül 2026
PDF İndir
EN TR

Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts

Öz

In this study, harmful content detection in Turkish texts is addressed as a three-class text classification problem comprising Harmless, Profanity and Sexist Discourse. For this purpose, open-access Turkish datasets and data compiled within the scope of the TÜBİTAK 1001 project were combined, and after removing cross-source duplicates and conflicting duplicates, a final dataset of 12051 texts was used. In the experimental process, traditional machine learning models and Transformer-based language models were compared under the same dataset, the same class structure and a 5-fold group-based stratified cross-validation protocol. For the traditional models, in addition to word- and character-level TF-IDF representations, Word2Vec- and FastText-based representations were used to evaluate Linear SVM, Logistic Regression, Decision Tree, Random Forest, Extra Trees and XGBoost algorithms. For the Transformer-based models, BERTurk, Electra-TR, ConvBERTurk and XLM-RoBERTa were adapted to the three-class classification task. Among the traditional models, the highest performance was obtained with Logistic Regression on TF-IDF representations (82.82% macro F1-score); Linear SVM produced very close results, while FastText- and Word2Vec-based representations remained comparatively lower. The obtained results showed that Transformer-based models offer higher and more balanced performance than traditional machine learning models. Among all models, the highest success was achieved with the ConvBERTurk model, which reached a 90.93% macro F1-score, 90.93% accuracy and a 98.09% ROC-AUC value. The BERTurk model exhibited a performance very close to ConvBERTurk with a 90.28% macro F1-score. A paired t-test over the five folds shows that this difference is not statistically significant (t(4) = 2.36, P=0.078). These findings indicate that contextual language representations are effective in Turkish harmful content classification. The study contributes to the literature by addressing Turkish harmful content detection as a multi-class problem that also includes the distinction between profanity and sexist discourse, and by comparing traditional models and Transformer-based models under the same leakage-controlled experimental protocol.

Anahtar Kelimeler

Etik Beyan

Ethics committee approval was not required for this study because of there was no study on animals or humans.

Teşekkür

This study was supported by the Scientific and Technological Research Council of Türkiye (TÜBİTAK) under the 1001 Program, Project No. 124E056. The authors would like to thank all members of the project team for their valuable contributions to the compilation and labeling of the dataset. During the preparation of this work, the authors used Claude, an AI language model developed by Anthropic, to assist in drafting and refining sections of the manuscript. After using this tool, the authors reviewed and edited the content to ensure its accuracy and appropriateness. The authors take full responsibility for the content of the published article.

Kaynakça

  1. Aksoy, Ç., Demirezen, M. U., & Sağıroğlu, Ş. (2026). Hate speech detection in Turkish: An ensemble transformer-based deep learning approach. Engineering Applications of Artificial Intelligence, 164, 113147. https://doi.org/10.1016/j.engappai.2025.113147
  2. Aliyeva, Ç. O., & Yağanoğlu, M. (2025). Deep learning approach to detect cyberbullying on Twitter. Multimedia Tools and Applications, 84(19), 20497–20520. https://doi.org/10.1007/s11042-024-19869-3
  3. Altinel, A. B., Karatas Baydogmus, G., Sahin, S., & Gurbuz, M. Z. (2024). So-haTRed: A novel hybrid system for Turkish hate speech detection in social media with ensemble deep learning improved by BERT and clustered-graph networks. IEEE Access, 12, 86252–86270. https://doi.org/10.1109/ACCESS.2024.3415350
  4. Barkhordar, E., Topçu, I. S., & Hürriyetoğlu, A. (2024, March). Team Curie at HSD-2Lang 2024: Hate speech detection in Turkish and Arabic tweets using BERT-based models. In Proceedings of the 7th Workshop on Challenges and Applications of Automated Extraction of Socio-Political Events from Text (CASE 2024) (pp. 215–220).
  5. Bojanowski, P., Grave, E., Joulin, A., & Mikolov, T. (2017). Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5, 135–146. https://doi.org/10.1162/tacl_a_00051
  6. Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324
  7. Breiman, L., Friedman, J. H., Olshen, R. A., & Stone, C. J. (1984). Classification and regression trees. Wadsworth International Group.
  8. Büyükdemirci, K., Kucukkaya, I. E., Ölmez, E., & Toraman, C. (2024). JL-Hate: An annotated dataset for joint learning of hate speech and target detection. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 9543–9553).

Ayrıntılar

Birincil Dil

İngilizce

Konular

Bilgi Sistemleri (Diğer)

Bölüm

Araştırma Makalesi

Yayımlanma Tarihi

15 Eylül 2026

Gönderilme Tarihi

17 Temmuz 2026

Kabul Tarihi

23 Ağustos 2026

Yayımlandığı Sayı

Yıl 2026 Cilt: 9 Sayı: 5

Kaynak Göster

APA
Karasekreter, N., Balım, C., & Aslan, Ö. (2026). Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts. Black Sea Journal of Engineering and Science, 9(5), 2676-2691. https://doi.org/10.34248/bsengineering.1996841
AMA
1.Karasekreter N, Balım C, Aslan Ö. Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts. BSJ Eng. Sci. 2026;9(5):2676-2691. doi:10.34248/bsengineering.1996841
Chicago
Karasekreter, Naim, Caner Balım, ve Özkan Aslan. 2026. “Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts”. Black Sea Journal of Engineering and Science 9 (5): 2676-91. https://doi.org/10.34248/bsengineering.1996841.
EndNote
Karasekreter N, Balım C, Aslan Ö (01 Eylül 2026) Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts. Black Sea Journal of Engineering and Science 9 5 2676–2691.
IEEE
[1]N. Karasekreter, C. Balım, ve Ö. Aslan, “Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts”, BSJ Eng. Sci., c. 9, sy 5, ss. 2676–2691, Eyl. 2026, doi: 10.34248/bsengineering.1996841.
ISNAD
Karasekreter, Naim - Balım, Caner - Aslan, Özkan. “Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts”. Black Sea Journal of Engineering and Science 9/5 (01 Eylül 2026): 2676-2691. https://doi.org/10.34248/bsengineering.1996841.
JAMA
1.Karasekreter N, Balım C, Aslan Ö. Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts. BSJ Eng. Sci. 2026;9:2676–2691.
MLA
Karasekreter, Naim, vd. “Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts”. Black Sea Journal of Engineering and Science, c. 9, sy 5, Eylül 2026, ss. 2676-91, doi:10.34248/bsengineering.1996841.
Vancouver
1.Naim Karasekreter, Caner Balım, Özkan Aslan. Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts. BSJ Eng. Sci. 01 Eylül 2026;9(5):2676-91. doi:10.34248/bsengineering.1996841