Research Article

Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts

Volume: 9 Number: 5 September 15, 2026
EN TR

Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts

Abstract

In this study, harmful content detection in Turkish texts is addressed as a three-class text classification problem comprising Harmless, Profanity and Sexist Discourse. For this purpose, open-access Turkish datasets and data compiled within the scope of the TÜBİTAK 1001 project were combined, and after removing cross-source duplicates and conflicting duplicates, a final dataset of 12051 texts was used. In the experimental process, traditional machine learning models and Transformer-based language models were compared under the same dataset, the same class structure and a 5-fold group-based stratified cross-validation protocol. For the traditional models, in addition to word- and character-level TF-IDF representations, Word2Vec- and FastText-based representations were used to evaluate Linear SVM, Logistic Regression, Decision Tree, Random Forest, Extra Trees and XGBoost algorithms. For the Transformer-based models, BERTurk, Electra-TR, ConvBERTurk and XLM-RoBERTa were adapted to the three-class classification task. Among the traditional models, the highest performance was obtained with Logistic Regression on TF-IDF representations (82.82% macro F1-score); Linear SVM produced very close results, while FastText- and Word2Vec-based representations remained comparatively lower. The obtained results showed that Transformer-based models offer higher and more balanced performance than traditional machine learning models. Among all models, the highest success was achieved with the ConvBERTurk model, which reached a 90.93% macro F1-score, 90.93% accuracy and a 98.09% ROC-AUC value. The BERTurk model exhibited a performance very close to ConvBERTurk with a 90.28% macro F1-score. A paired t-test over the five folds shows that this difference is not statistically significant (t(4) = 2.36, P=0.078). These findings indicate that contextual language representations are effective in Turkish harmful content classification. The study contributes to the literature by addressing Turkish harmful content detection as a multi-class problem that also includes the distinction between profanity and sexist discourse, and by comparing traditional models and Transformer-based models under the same leakage-controlled experimental protocol.

Keywords

Ethical Statement

Ethics committee approval was not required for this study because of there was no study on animals or humans.

Thanks

This study was supported by the Scientific and Technological Research Council of Türkiye (TÜBİTAK) under the 1001 Program, Project No. 124E056. The authors would like to thank all members of the project team for their valuable contributions to the compilation and labeling of the dataset. During the preparation of this work, the authors used Claude, an AI language model developed by Anthropic, to assist in drafting and refining sections of the manuscript. After using this tool, the authors reviewed and edited the content to ensure its accuracy and appropriateness. The authors take full responsibility for the content of the published article.

References

  1. Aksoy, Ç., Demirezen, M. U., & Sağıroğlu, Ş. (2026). Hate speech detection in Turkish: An ensemble transformer-based deep learning approach. Engineering Applications of Artificial Intelligence, 164, 113147. https://doi.org/10.1016/j.engappai.2025.113147
  2. Aliyeva, Ç. O., & Yağanoğlu, M. (2025). Deep learning approach to detect cyberbullying on Twitter. Multimedia Tools and Applications, 84(19), 20497–20520. https://doi.org/10.1007/s11042-024-19869-3
  3. Altinel, A. B., Karatas Baydogmus, G., Sahin, S., & Gurbuz, M. Z. (2024). So-haTRed: A novel hybrid system for Turkish hate speech detection in social media with ensemble deep learning improved by BERT and clustered-graph networks. IEEE Access, 12, 86252–86270. https://doi.org/10.1109/ACCESS.2024.3415350
  4. Barkhordar, E., Topçu, I. S., & Hürriyetoğlu, A. (2024, March). Team Curie at HSD-2Lang 2024: Hate speech detection in Turkish and Arabic tweets using BERT-based models. In Proceedings of the 7th Workshop on Challenges and Applications of Automated Extraction of Socio-Political Events from Text (CASE 2024) (pp. 215–220).
  5. Bojanowski, P., Grave, E., Joulin, A., & Mikolov, T. (2017). Enriching word vectors with subword information. Transactions of the Association for Computational Linguistics, 5, 135–146. https://doi.org/10.1162/tacl_a_00051
  6. Breiman, L. (2001). Random forests. Machine Learning, 45(1), 5–32. https://doi.org/10.1023/A:1010933404324
  7. Breiman, L., Friedman, J. H., Olshen, R. A., & Stone, C. J. (1984). Classification and regression trees. Wadsworth International Group.
  8. Büyükdemirci, K., Kucukkaya, I. E., Ölmez, E., & Toraman, C. (2024). JL-Hate: An annotated dataset for joint learning of hate speech and target detection. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 9543–9553).

Details

Primary Language

English

Subjects

Information Systems (Other)

Journal Section

Research Article

Publication Date

September 15, 2026

Submission Date

July 17, 2026

Acceptance Date

August 23, 2026

Published in Issue

Year 2026 Volume: 9 Number: 5

APA
Karasekreter, N., Balım, C., & Aslan, Ö. (2026). Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts. Black Sea Journal of Engineering and Science, 9(5), 2676-2691. https://doi.org/10.34248/bsengineering.1996841
AMA
1.Karasekreter N, Balım C, Aslan Ö. Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts. BSJ Eng. Sci. 2026;9(5):2676-2691. doi:10.34248/bsengineering.1996841
Chicago
Karasekreter, Naim, Caner Balım, and Özkan Aslan. 2026. “Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts”. Black Sea Journal of Engineering and Science 9 (5): 2676-91. https://doi.org/10.34248/bsengineering.1996841.
EndNote
Karasekreter N, Balım C, Aslan Ö (September 1, 2026) Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts. Black Sea Journal of Engineering and Science 9 5 2676–2691.
IEEE
[1]N. Karasekreter, C. Balım, and Ö. Aslan, “Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts”, BSJ Eng. Sci., vol. 9, no. 5, pp. 2676–2691, Sept. 2026, doi: 10.34248/bsengineering.1996841.
ISNAD
Karasekreter, Naim - Balım, Caner - Aslan, Özkan. “Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts”. Black Sea Journal of Engineering and Science 9/5 (September 1, 2026): 2676-2691. https://doi.org/10.34248/bsengineering.1996841.
JAMA
1.Karasekreter N, Balım C, Aslan Ö. Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts. BSJ Eng. Sci. 2026;9:2676–2691.
MLA
Karasekreter, Naim, et al. “Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts”. Black Sea Journal of Engineering and Science, vol. 9, no. 5, Sept. 2026, pp. 2676-91, doi:10.34248/bsengineering.1996841.
Vancouver
1.Naim Karasekreter, Caner Balım, Özkan Aslan. Machine Learning and Deep Learning Based Classification of Profanity and Sexist Discourse in Turkish Texts. BSJ Eng. Sci. 2026 Sep. 1;9(5):2676-91. doi:10.34248/bsengineering.1996841