A Deep Learning based Approach for Identifying Spoken languages in Turkey
Abstract
Spoken language identification (SLID) systems are used in areas such as human–machine interaction and security. In forensic and intelligence work, identifying the language in short or degraded recordings is important for criminal, immigration, and national security purposes. To address the specific challenges of these domains, this study focuses on extracting robust acoustic features that remain discriminative even in noisy environments and with limited temporal data. In this study, a speech dataset was constructed using languages that are widely spoken in Türkiye, including those used by foreign residents and visitors. The recordings span six languages: Turkish, Kurdish, Arabic, Persian, Russian, and German. The task of recognizing the spoken language was formulated as a multi-class classification problem. To address this task, several state-of-the-art deep learning architectures were employed, including Convolutional Neural Networks (CNNs), Transformer-based encoders, Bidirectional Long Short-Term Memory (BiLSTM) networks, and Conformer encoder models, which integrate convolutional operations with self-attention mechanisms to effectively process short utterances and complex signal patterns. These models were trained and evaluated under different data configurations to assess their robustness and generalization. Model performance was evaluated using widely accepted metrics such as accuracy, recall, and macro-averaged F1-score, with a particular focus on handling unbalanced data and phonetically similar language pairs. Experimental results reveal that the Conformer encoder demonstrated superior robustness, achieving the highest accuracy of 99.89% and a macro-averaged F1-score of 0.9914 on the population-weighted unbalanced dataset. In contrast, the Transformer model showed significant sensitivity to class imbalance, where its macro F1-score dropped to 0.8179.
Keywords
References
- Göç İdaresi Başkanlığı, T.C. İçişleri Bakanlığı, "T.C. İçişleri Bakanlığı Göç İdaresi Başkanlığı", https://www.goc.gov.tr/ , Access date: 18.05.2025.
- Veri Portalı, TÜİK, "TÜİK - Veri Portalı", https://data.tuik.gov.tr/Kategori/GetKategori?p=Nufus-ve-Demografi-109 , Access date: 18.05.2025.
- Marketing Türkiye, "Türkiye’ye sekiz ayda 40 milyon turist geldi: Rusya ve Almanya, başı çekiyor", https://www.marketingturkiye.com.tr/haberler/turkiyeye-sekiz-ayda-40-milyon-turist-geldi-rusya-ve-almanya-basi-cekiyor/ , Access date: 31.01.2026.
- Gonzalez-Dominguez, J., Lopez-Moreno, I., Moreno, P. J., and Gonzalez-Rodriguez, J., "Frame-by-frame language identification in short utterances using deep neural networks", Neural Networks, 64: 49–58, (2015). DOI: https://doi.org/10.1016/j.neunet.2014.08.006
- Li, H., Ma, B., and Lee, K.A., "Spoken language recognition: from fundamentals to practice", Proceedings of the IEEE, 101(5): 1136-1159, (2013). DOI: https://doi.org/10.1109/JPROC.2012.2237151
- O’Shaughnessy, D., "Spoken language identification: An overview of past and present research trends", Speech Communication, 167: 103167, (2025). DOI: https://doi.org/10.1016/j.specom.2024.103167
- Anidjar, O. H. and Yozevitch, R., "Enhancing neural spoken language recognition: An exploration with multilingual datasets", arXiv preprint, arXiv:2501.11065, (2025). DOI: https://doi.org/10.48550/arXiv.2501.11065
- Çelik, Y., "Application of deep learning for voice command classification in Turkish language", Bitlis Eren Üniversitesi Fen Bilimleri Dergisi, 13(3): 701–708, (2024). DOI: https://doi.org/10.17798/bitlisfen.1477191
Details
Primary Language
English
Subjects
Natural Language Processing
Journal Section
Research Article
Early Pub Date
July 24, 2026
Publication Date
-
Submission Date
June 22, 2025
Acceptance Date
May 14, 2026
Published in Issue
Year 2026 Number: Advanced Online Publication