Automatic Speech Recognition from Throat Microphone Signals of the NATO Phonetic Alphabet Spoken by Turkish Speakers: Evaluation of Classical and Deep Techniques under Push-to-Talk Conditions
Öz
Anahtar Kelimeler
Automatic speech recognition (ASR), Machine learning, NATO phonetic alphabet, Push-to-talk (PTT), Throat microphone
Destekleyen Kurum
Etik Beyan
Kaynakça
- Y. Görmez, “Customized deep learning based Turkish automatic speech recognition system supported by language model,” PeerJ. Computer Science, vol. 10, p. e1981, 2024.
- Z. Y. Ren, Nurmement, H. Wang, and W. Slamu, “Exploring Turkish speech recognition via hybrid CTC/ attention architecture and multi-feature fusion network,” arXiv preprint arXiv:2308.10654 [cs.SD], 2023. (For arXiv preprints, it’s good practice to include the arXiv ID).
- E. T. Erzin and T. Tugtekin, “Improving phoneme recognition of throat microphone speech recordings using transfer learning,” Speech Communication, vol. 129, p. 104764, 2021. (Added article number/page if available, assuming “10” was part of it. If it was just the number of the article, no. 10 is fine).
- J. L. K. E. T. Fendji, D. C. M. Diane, B. O. Yenke, and M. Atemkeng, “Automatic speech recognition using limited vocabulary: A survey,” Applied Artificial Intelligence: AAI, vol. 36, no. 1, 2022.
- NATO, Allied Communications Publication ACP 125(G) Communication Instructions Radiotelephone Proce- dures. 2008. (This is a report/publication from an organization).
- A. Veysov, “silero-vad: Silero VAD: pre-trained enterprise-grade Voice Activity Detector,” GitHub, 2022. [Online]. Available: https:/ /github.com/snakers4/silero-vad (For software/code, it’s best to treat it like a technical report or use the most formal citation possible, often including the repository host and year. The original example didn’t have a year or formal publisher).
- T. Ko, V. Peddinti, D. Povey, and S. Khudanpur, “Audio augmentation for speech recognition,” in Proc. In - terspeech 2015, Dresden, Germany, 2015, pp. 3586-3589. (For conference papers, it’s good to include the location of the conference if known).
- D. S. C. Park, W. William, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A Simple Data Augmentation Method for Automatic Speech Recognition,” arXiv preprint arXiv:1904.08779 [eess.AS], 2019. (Added arXiv ID).
- C. Shorten and T. M. K. Taghi, “A survey on image data augmentation for deep learning,” Journal of Big Data, vol. 6, no. 1, 2019.
- S. M. Davis and P. Mermelstein, “Comparison of parametric representations for monosyllabic word recog- nition in continuously spoken sentences,” IEEE Transactions on Acoustics, Speech, and Signal Processing, vol. 28, no. 4, pp. 357-366, 1980. (Corrected author ‘P’ to ‘P. Mermelstein’ as per common authorship in this field, assuming this is the full name, and added volume, number, and pages).