Research Article

Hybrid Model of Attention-Guided Deep Learning and Metaheuristic Optimization for Speech Emotion Recognition

Volume: 17 Number: 2 July 27, 2026
TR EN

Hybrid Model of Attention-Guided Deep Learning and Metaheuristic Optimization for Speech Emotion Recognition

Abstract

Speech Emotion Recognition (SER) technologies, which have become a significant part of the human-computer interaction field, are artificial intelligence-based systems that automatically detect emotional states in human speech signals. The working principle of these systems is the process of matching acoustic features extracted from voice signals with emotional labels through machine learning. Despite significant advancements, current SER technologies still face some challenges. Chief among these is the inadequate representation of local and global patterns. This study presents an innovative hybrid model to overcome these challenges in emotion representation. The model consists of two phases. The first phase represents the basic model, which involves an end-to-end learning process with the extraction of feature representations of emotions and their matching with emotion labels. The second phase is the process where feature optimization takes place. The first phase consists of steps including local feature extraction, channel-level feature activation, temporal dependency representation, and finally, focusing on key time periods. The second phase is the process of optimizing features using eighteen different optimization configurations derived from the integration of six meta-heuristic algorithms with three machine learning paradigms. Experiments were conducted on the EMO-DB dataset. In a 5-fold cross-validation method, the proposed hybrid model achieved an unweighted accuracy (UA) of 83.73%, demonstrating a 7.8% increase in unweighted accuracy compared to the base model, while reducing feature dimensionality by 54.8%. These results confirm that the proposed hybrid model provides high accuracy and efficient feature representations.

Keywords

References

  1. [1] J. Ye et al., “Multi-modal depression detection based on emotional audio and evaluation text”, J. Affect. Disord., vol. 295, pp. 904-913, 2021.
  2. [2] F. Yin, J. Du, X. Xu, and L. Zhao, “Depression Detection in Speech Using Transformer and Parallel Convolutional Neural Networks”, Electronics 2023, vol. 12, no. 2, p. 328, 2023.
  3. [3] E. W. McGinnis et al., “Giving Voice to Vulnerable Children: Machine Learning Analysis of Speech Detects Anxiety and Depression in Early Childhood”, IEEE J. Biomed. Health Inform., vol. 23, no. 6, pp. 2294-2301, 2019.
  4. [4] W. Xian, “Speech Emotion Recognition Application for Education”, BCP Education & Psychology, vol. 7, pp. 378-383, 2022.
  5. [5] M. Płaza et al., “Emotion Recognition Method for Call/Contact Centre Systems”, Applied Sciences 2022, vol. 12, no. 21, p. 10951, 2022.
  6. [6] Y. Feng and L. Devillers, “End-to-End Continuous Speech Emotion Recognition in Real-life Customer Service Call Center Conversations”, 2023 11th International Conference on Affective Computing and Intelligent Interaction Workshops and Demos, ACIIW 2023, 2023.
  7. [7] I. Gurowiec and N. Nissim, “Speech emotion recognition systems and their security aspects”, Artif. Intell. Rev., vol. 57, no. 6, pp. 1-45, 2024.
  8. [8] Z. Jia, J. Yu, H. Long, and D. Tang, “Coverage-Guaranteed Speech Emotion Recognition via Calibrated Uncertainty-Adaptive Prediction Sets”, Accessed: Jul. 29, 2025. [Online]. Available: https://arxiv.org/abs/2503.22712v4

Details

Primary Language

English

Subjects

Audio Processing, Human-Computer Interaction, Neural Networks, Artificial Intelligence (Other)

Journal Section

Research Article

Publication Date

July 27, 2026

Submission Date

February 20, 2026

Acceptance Date

March 26, 2026

Published in Issue

Year 2026 Volume: 17 Number: 2

APA
Donuk, K. (2026). Hybrid Model of Attention-Guided Deep Learning and Metaheuristic Optimization for Speech Emotion Recognition. Dicle Üniversitesi Mühendislik Fakültesi Mühendislik Dergisi, 17(2). https://doi.org/10.24012/dumf.1893980
AMA
1.Donuk K. Hybrid Model of Attention-Guided Deep Learning and Metaheuristic Optimization for Speech Emotion Recognition. DUJE. 2026;17(2). doi:10.24012/dumf.1893980
Chicago
Donuk, Kenan. 2026. “Hybrid Model of Attention-Guided Deep Learning and Metaheuristic Optimization for Speech Emotion Recognition”. Dicle Üniversitesi Mühendislik Fakültesi Mühendislik Dergisi 17 (2). https://doi.org/10.24012/dumf.1893980.
EndNote
Donuk K (July 1, 2026) Hybrid Model of Attention-Guided Deep Learning and Metaheuristic Optimization for Speech Emotion Recognition. Dicle Üniversitesi Mühendislik Fakültesi Mühendislik Dergisi 17 2
IEEE
[1]K. Donuk, “Hybrid Model of Attention-Guided Deep Learning and Metaheuristic Optimization for Speech Emotion Recognition”, DUJE, vol. 17, no. 2, July 2026, doi: 10.24012/dumf.1893980.
ISNAD
Donuk, Kenan. “Hybrid Model of Attention-Guided Deep Learning and Metaheuristic Optimization for Speech Emotion Recognition”. Dicle Üniversitesi Mühendislik Fakültesi Mühendislik Dergisi 17/2 (July 1, 2026). https://doi.org/10.24012/dumf.1893980.
JAMA
1.Donuk K. Hybrid Model of Attention-Guided Deep Learning and Metaheuristic Optimization for Speech Emotion Recognition. DUJE. 2026;17. doi:10.24012/dumf.1893980.
MLA
Donuk, Kenan. “Hybrid Model of Attention-Guided Deep Learning and Metaheuristic Optimization for Speech Emotion Recognition”. Dicle Üniversitesi Mühendislik Fakültesi Mühendislik Dergisi, vol. 17, no. 2, July 2026, doi:10.24012/dumf.1893980.
Vancouver
1.Kenan Donuk. Hybrid Model of Attention-Guided Deep Learning and Metaheuristic Optimization for Speech Emotion Recognition. DUJE. 2026 Jul. 1;17(2). doi:10.24012/dumf.1893980