Research Article

A Deep Learning based Approach for Identifying Spoken languages in Turkey

Number: Advanced Online Publication Early Pub Date: July 24, 2026
EN

A Deep Learning based Approach for Identifying Spoken languages in Turkey

Abstract

Spoken language identification (SLID) systems are used in areas such as human–machine interaction and security. In forensic and intelligence work, identifying the language in short or degraded recordings is important for criminal, immigration, and national security purposes. To address the specific challenges of these domains, this study focuses on extracting robust acoustic features that remain discriminative even in noisy environments and with limited temporal data. In this study, a speech dataset was constructed using languages that are widely spoken in Türkiye, including those used by foreign residents and visitors. The recordings span six languages: Turkish, Kurdish, Arabic, Persian, Russian, and German. The task of recognizing the spoken language was formulated as a multi-class classification problem. To address this task, several state-of-the-art deep learning architectures were employed, including Convolutional Neural Networks (CNNs), Transformer-based encoders, Bidirectional Long Short-Term Memory (BiLSTM) networks, and Conformer encoder models, which integrate convolutional operations with self-attention mechanisms to effectively process short utterances and complex signal patterns. These models were trained and evaluated under different data configurations to assess their robustness and generalization. Model performance was evaluated using widely accepted metrics such as accuracy, recall, and macro-averaged F1-score, with a particular focus on handling unbalanced data and phonetically similar language pairs. Experimental results reveal that the Conformer encoder demonstrated superior robustness, achieving the highest accuracy of 99.89% and a macro-averaged F1-score of 0.9914 on the population-weighted unbalanced dataset. In contrast, the Transformer model showed significant sensitivity to class imbalance, where its macro F1-score dropped to 0.8179.

Keywords

References

  1. Göç İdaresi Başkanlığı, T.C. İçişleri Bakanlığı, "T.C. İçişleri Bakanlığı Göç İdaresi Başkanlığı", https://www.goc.gov.tr/ , Access date: 18.05.2025.
  2. Veri Portalı, TÜİK, "TÜİK - Veri Portalı", https://data.tuik.gov.tr/Kategori/GetKategori?p=Nufus-ve-Demografi-109 , Access date: 18.05.2025.
  3. Marketing Türkiye, "Türkiye’ye sekiz ayda 40 milyon turist geldi: Rusya ve Almanya, başı çekiyor", https://www.marketingturkiye.com.tr/haberler/turkiyeye-sekiz-ayda-40-milyon-turist-geldi-rusya-ve-almanya-basi-cekiyor/ , Access date: 31.01.2026.
  4. Gonzalez-Dominguez, J., Lopez-Moreno, I., Moreno, P. J., and Gonzalez-Rodriguez, J., "Frame-by-frame language identification in short utterances using deep neural networks", Neural Networks, 64: 49–58, (2015). DOI: https://doi.org/10.1016/j.neunet.2014.08.006
  5. Li, H., Ma, B., and Lee, K.A., "Spoken language recognition: from fundamentals to practice", Proceedings of the IEEE, 101(5): 1136-1159, (2013). DOI: https://doi.org/10.1109/JPROC.2012.2237151
  6. O’Shaughnessy, D., "Spoken language identification: An overview of past and present research trends", Speech Communication, 167: 103167, (2025). DOI: https://doi.org/10.1016/j.specom.2024.103167
  7. Anidjar, O. H. and Yozevitch, R., "Enhancing neural spoken language recognition: An exploration with multilingual datasets", arXiv preprint, arXiv:2501.11065, (2025). DOI: https://doi.org/10.48550/arXiv.2501.11065
  8. Çelik, Y., "Application of deep learning for voice command classification in Turkish language", Bitlis Eren Üniversitesi Fen Bilimleri Dergisi, 13(3): 701–708, (2024). DOI: https://doi.org/10.17798/bitlisfen.1477191

Details

Primary Language

English

Subjects

Natural Language Processing

Journal Section

Research Article

Early Pub Date

July 24, 2026

Publication Date

-

Submission Date

June 22, 2025

Acceptance Date

May 14, 2026

Published in Issue

Year 2026 Number: Advanced Online Publication

APA
Ertekin, M., & Erdem, H. (2026). A Deep Learning based Approach for Identifying Spoken languages in Turkey. Gazi University Journal of Science, Advanced Online Publication. https://doi.org/10.35378/gujs.1724816
AMA
1.Ertekin M, Erdem H. A Deep Learning based Approach for Identifying Spoken languages in Turkey. Gazi University Journal of Science. 2026;(Advanced Online Publication). doi:10.35378/gujs.1724816
Chicago
Ertekin, Müge, and Hamit Erdem. 2026. “A Deep Learning Based Approach for Identifying Spoken Languages in Turkey”. Gazi University Journal of Science, no. Advanced Online Publication. https://doi.org/10.35378/gujs.1724816.
EndNote
Ertekin M, Erdem H (July 1, 2026) A Deep Learning based Approach for Identifying Spoken languages in Turkey. Gazi University Journal of Science Advanced Online Publication
IEEE
[1]M. Ertekin and H. Erdem, “A Deep Learning based Approach for Identifying Spoken languages in Turkey”, Gazi University Journal of Science, no. Advanced Online Publication, July 2026, doi: 10.35378/gujs.1724816.
ISNAD
Ertekin, Müge - Erdem, Hamit. “A Deep Learning Based Approach for Identifying Spoken Languages in Turkey”. Gazi University Journal of Science. Advanced Online Publication (July 1, 2026). https://doi.org/10.35378/gujs.1724816.
JAMA
1.Ertekin M, Erdem H. A Deep Learning based Approach for Identifying Spoken languages in Turkey. Gazi University Journal of Science. 2026. doi:10.35378/gujs.1724816.
MLA
Ertekin, Müge, and Hamit Erdem. “A Deep Learning Based Approach for Identifying Spoken Languages in Turkey”. Gazi University Journal of Science, no. Advanced Online Publication, July 2026, doi:10.35378/gujs.1724816.
Vancouver
1.Müge Ertekin, Hamit Erdem. A Deep Learning based Approach for Identifying Spoken languages in Turkey. Gazi University Journal of Science. 2026 Jul. 1;(Advanced Online Publication). doi:10.35378/gujs.1724816