Research Article

Automatic Speech Recognition from Throat Microphone Signals of the NATO Phonetic Alphabet Spoken by Turkish Speakers: Evaluation of Classical and Deep Techniques under Push-to-Talk Conditions

Volume: 5 Number: 2 November 13, 2025
TR EN

Automatic Speech Recognition from Throat Microphone Signals of the NATO Phonetic Alphabet Spoken by Turkish Speakers: Evaluation of Classical and Deep Techniques under Push-to-Talk Conditions

Abstract

This study evaluates the performance of various Automatic Speech Recognition (ASR) techniques applied exclusively to recordings captured using throat microphones (TM). The objective is to explore their applicability in Push-to-Talk (PTT) operational conditions, where traditional air microphones are limited by environmental noise. A corpus was constructed with 10 native Turkish speakers enunciating the NATO phonetic alphabet. Signals were segmented using Silero VAD, resampled to 16 kHz, and augmented to robustify the models against noise and variations. Two feature extraction approaches were employed: Mel-frequency cepstral coefficients (MFCCs) and Wav2Vec2 embeddings reduced by Principal Component Analysis (PCA). Subsequently, five supervised classifiers were trained and compared: SVM, RF, KNN, MLP, and LightGBM. Evaluation metrics included overall accuracy and Word Error Rate (WER). Results demonstrate the technical feasibility of ASR with laryngeal signals, identifying the combination of LightGBM with MFCCs as the most robust (86.38% accuracy, 0.000 WER) and confirming the potential of RF with MFCCs (84.62% accuracy, 0.000 WER). This work establishes an experimental foundation for the development of robust and low-cost ASR systems in noisy environments. In this context, throat microphones offer a crucial alternative.

Keywords

Automatic speech recognition (ASR), Machine learning, NATO phonetic alphabet, Push-to-talk (PTT), Throat microphone

Supporting Institution

This research did not receive any specific grant from funding agencies in the public, commercial, or not-for-profit sectors.

Ethical Statement

This study was conducted with the approval of the Ethics Committee for Social and Human Sciences Research of Ondokuz Mayıs University. The approval was granted on 29 November 2024 under the decision number 2024-1118. The research involved voice recordings and interviews carried out as part of a master’s thesis titled “Machine Learning-Based Analysis and Resolution of Multilingual Pronunciation Issues in the NATO Phonetic Alphabet”, under the supervision of Dr. Öğr. Üyesi Selim Aras.

References

  1. Y. Görmez, “Customized deep learning based Turkish automatic speech recognition system supported by language model,” PeerJ. Computer Science, vol. 10, p. e1981, 2024.
  2. Z. Y. Ren, Nurmement, H. Wang, and W. Slamu, “Exploring Turkish speech recognition via hybrid CTC/ attention architecture and multi-feature fusion network,” arXiv preprint arXiv:2308.10654 [cs.SD], 2023. (For arXiv preprints, it’s good practice to include the arXiv ID).
  3. E. T. Erzin and T. Tugtekin, “Improving phoneme recognition of throat microphone speech recordings using transfer learning,” Speech Communication, vol. 129, p. 104764, 2021. (Added article number/page if available, assuming “10” was part of it. If it was just the number of the article, no. 10 is fine).
  4. J. L. K. E. T. Fendji, D. C. M. Diane, B. O. Yenke, and M. Atemkeng, “Automatic speech recognition using limited vocabulary: A survey,” Applied Artificial Intelligence: AAI, vol. 36, no. 1, 2022.
  5. NATO, Allied Communications Publication ACP 125(G) Communication Instructions Radiotelephone Proce- dures. 2008. (This is a report/publication from an organization).
  6. A. Veysov, “silero-vad: Silero VAD: pre-trained enterprise-grade Voice Activity Detector,” GitHub, 2022. [Online]. Available: https:/ /github.com/snakers4/silero-vad (For software/code, it’s best to treat it like a technical report or use the most formal citation possible, often including the repository host and year. The original example didn’t have a year or formal publisher).
  7. T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio augmentation for speech recognition,” in Proc. In - terspeech 2015, Dresden, Germany, 2015, pp. 3586-3589. (For conference papers, it’s good to include the location of the conference if known).
  8. D. S. C. Park, W. William, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition,” arXiv preprint arXiv:1904.08779 [eess.AS], 2019. (Added arXiv ID).
  9. C. Shorten and T. M. K. Taghi, “A survey on image data augmentation for deep learning,” Journal of Big Data, vol. 6, no. 1, 2019.
  10. S. M. Davis and P. Mermelstein, “Comparison of parametric representations for monosyllabic word recog- nition in continuously spoken sentences,” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 28, no. 4, pp. 357-366, 1980. (Corrected author ‘P’ to ‘P. Mermelstein’ as per common authorship in this field, assuming this is the full name, and added volume, number, and pages).
APA
Velazquez, J., & Aras, S. (2025). Automatic Speech Recognition from Throat Microphone Signals of the NATO Phonetic Alphabet Spoken by Turkish Speakers: Evaluation of Classical and Deep Techniques under Push-to-Talk Conditions. OMÜ Mühendislik Bilimleri Ve Teknolojisi Dergisi, 5(2), 73-91. https://izlik.org/JA46SH85WL
AMA
1.Velazquez J, Aras S. Automatic Speech Recognition from Throat Microphone Signals of the NATO Phonetic Alphabet Spoken by Turkish Speakers: Evaluation of Classical and Deep Techniques under Push-to-Talk Conditions. OMUJEST. 2025;5(2):73-91. https://izlik.org/JA46SH85WL
Chicago
Velazquez, Julio, and Selim Aras. 2025. “Automatic Speech Recognition from Throat Microphone Signals of the NATO Phonetic Alphabet Spoken by Turkish Speakers: Evaluation of Classical and Deep Techniques under Push-to-Talk Conditions”. OMÜ Mühendislik Bilimleri Ve Teknolojisi Dergisi 5 (2): 73-91. https://izlik.org/JA46SH85WL.
EndNote
Velazquez J, Aras S (November 1, 2025) Automatic Speech Recognition from Throat Microphone Signals of the NATO Phonetic Alphabet Spoken by Turkish Speakers: Evaluation of Classical and Deep Techniques under Push-to-Talk Conditions. OMÜ Mühendislik Bilimleri ve Teknolojisi Dergisi 5 2 73–91.
IEEE
[1]J. Velazquez and S. Aras, “Automatic Speech Recognition from Throat Microphone Signals of the NATO Phonetic Alphabet Spoken by Turkish Speakers: Evaluation of Classical and Deep Techniques under Push-to-Talk Conditions”, OMUJEST, vol. 5, no. 2, pp. 73–91, Nov. 2025, [Online]. Available: https://izlik.org/JA46SH85WL
ISNAD
Velazquez, Julio - Aras, Selim. “Automatic Speech Recognition from Throat Microphone Signals of the NATO Phonetic Alphabet Spoken by Turkish Speakers: Evaluation of Classical and Deep Techniques under Push-to-Talk Conditions”. OMÜ Mühendislik Bilimleri ve Teknolojisi Dergisi 5/2 (November 1, 2025): 73-91. https://izlik.org/JA46SH85WL.
JAMA
1.Velazquez J, Aras S. Automatic Speech Recognition from Throat Microphone Signals of the NATO Phonetic Alphabet Spoken by Turkish Speakers: Evaluation of Classical and Deep Techniques under Push-to-Talk Conditions. OMUJEST. 2025;5:73–91.
MLA
Velazquez, Julio, and Selim Aras. “Automatic Speech Recognition from Throat Microphone Signals of the NATO Phonetic Alphabet Spoken by Turkish Speakers: Evaluation of Classical and Deep Techniques under Push-to-Talk Conditions”. OMÜ Mühendislik Bilimleri Ve Teknolojisi Dergisi, vol. 5, no. 2, Nov. 2025, pp. 73-91, https://izlik.org/JA46SH85WL.
Vancouver
1.Julio Velazquez, Selim Aras. Automatic Speech Recognition from Throat Microphone Signals of the NATO Phonetic Alphabet Spoken by Turkish Speakers: Evaluation of Classical and Deep Techniques under Push-to-Talk Conditions. OMUJEST [Internet]. 2025 Nov. 1;5(2):73-91. Available from: https://izlik.org/JA46SH85WL