Research Article

MFCCFusionnet: A CNN–DNN feature-fusion model for speech emotion recognition using MFCC representations

Volume: 21 Number: 73 July 20, 2026
TR EN

MFCCFusionnet: A CNN–DNN feature-fusion model for speech emotion recognition using MFCC representations

Abstract

Speech emotion recognition has become an increasingly important research area because of its potential to enhance human–machine interaction and support the development of emotion-aware systems. However, further comparative evidence is still needed on how different deep learning architectures perform when they are trained using the same acoustic feature representation. In this study, an experimental comparison was conducted using the RAVDESS (Ryerson Audio-Visual Database of Emotional Speech and Song) dataset. Mel-frequency cepstral coefficients (MFCCs) were used as the sole feature extraction method, and three deep learning architectures were designed and evaluated: a Convolutional Neural Network (CNN), a Deep Neural Network (DNN), and a hybrid feature-fusion model named MFCCFusionNet. The CNN and DNN models achieved classification accuracies of 81% and 76%, respectively. The proposed MFCCFusionNet model combined the learned representations obtained from the CNN and DNN sub-models and achieved the highest accuracy among the evaluated models, with a test accuracy of 83%. The findings indicate that MFCC-based representations can be effectively processed by different deep learning architectures and that feature-level fusion may improve classification performance in speech-based emotion recognition. The study provides preliminary evidence that feature-level fusion of CNN- and DNN-based MFCC representations may improve classification performance in speech emotion recognition.

Keywords

References

  1. Zhao, J., Mao, X., & Chen, L. (2019). Speech emotion recognition using deep 1D & 2D CNN LSTM networks. Biomedical Signal Processing and Control, 47, 312-323.
  2. Makhmudov, F., Kutlimuratov, A., & Cho, Y. I. (2024). Hybrid LSTM– Attention and CNN Model for Enhanced Speech Emotion Recognition. Applied Sciences, 14(23), 11342.
  3. Dinesh, P., Mishra, S. P., Warule, P., & Deb, S. (2024, December). Speech Emotion Recognition with DNN and Combination of CNN-LSTM. In TENCON 2024-2024 IEEE Region 10 Conference (TENCON) (pp. 1159-1162). IEEE.
  4. Yang, Z., Li, Z., Zhou, S., Zhang, L., & Serikawa, S. (2024). Speech emotion recognition based on multi-feature speed rate and LSTM. Neurocomputing, 601, 128177.
  5. Shetty, K. J., Shetty, S., & Shetty, M. (2024, April). Speech emotion recognition using lstm. In 2024 Third International Conference on Distributed Computing and Electrical Circuits and Electronics (ICDCECE) (pp. 1-6). IEEE.
  6. Şeker, A., Diri, B., Balık, H. H. (2017). Derin öğrenme yöntemleri ve uygulamaları hakkında bir inceleme. Gazi Mühendislik Bilimleri Dergisi, 3(3), 47–64.
  7. Ezz-Eldin, M., Khalaf, A. A. M., Hamed, H. F. A., & Hussein, A. I. (2021). Efficient feature-aware hybrid model of deep learning architectures for speech emotion recognition. IEEE Access, 9, 19999– 20011.
  8. Jakubec, M., Lieskovska, E., Jarina, R., Spisiak, M., & Kasak, P. (2024). Speech emotion recognition using transfer learning: Integration of advanced speaker embeddings and image recognition models. Applied Sciences, 14(21), 9981.

Details

Primary Language

English

Subjects

Information Systems (Other)

Journal Section

Research Article

Publication Date

July 20, 2026

Submission Date

April 29, 2026

Acceptance Date

June 29, 2026

Published in Issue

Year 2026 Volume: 21 Number: 73

APA
Gündoğan, N., & İşler, B. (2026). MFCCFusionnet: A CNN–DNN feature-fusion model for speech emotion recognition using MFCC representations. Anadolu Bil Meslek Yüksekokulu Dergisi, 21(73), 171-195. https://izlik.org/JA86UN57UM
AMA
1.Gündoğan N, İşler B. MFCCFusionnet: A CNN–DNN feature-fusion model for speech emotion recognition using MFCC representations. ABMYO Dergisi. 2026;21(73):171-195. https://izlik.org/JA86UN57UM
Chicago
Gündoğan, Neslihan, and Buket İşler. 2026. “MFCCFusionnet: A CNN–DNN Feature-Fusion Model for Speech Emotion Recognition Using MFCC Representations”. Anadolu Bil Meslek Yüksekokulu Dergisi 21 (73): 171-95. https://izlik.org/JA86UN57UM.
EndNote
Gündoğan N, İşler B (July 1, 2026) MFCCFusionnet: A CNN–DNN feature-fusion model for speech emotion recognition using MFCC representations. Anadolu Bil Meslek Yüksekokulu Dergisi 21 73 171–195.
IEEE
[1]N. Gündoğan and B. İşler, “MFCCFusionnet: A CNN–DNN feature-fusion model for speech emotion recognition using MFCC representations”, ABMYO Dergisi, vol. 21, no. 73, pp. 171–195, July 2026, [Online]. Available: https://izlik.org/JA86UN57UM
ISNAD
Gündoğan, Neslihan - İşler, Buket. “MFCCFusionnet: A CNN–DNN Feature-Fusion Model for Speech Emotion Recognition Using MFCC Representations”. Anadolu Bil Meslek Yüksekokulu Dergisi 21/73 (July 1, 2026): 171-195. https://izlik.org/JA86UN57UM.
JAMA
1.Gündoğan N, İşler B. MFCCFusionnet: A CNN–DNN feature-fusion model for speech emotion recognition using MFCC representations. ABMYO Dergisi. 2026;21:171–195.
MLA
Gündoğan, Neslihan, and Buket İşler. “MFCCFusionnet: A CNN–DNN Feature-Fusion Model for Speech Emotion Recognition Using MFCC Representations”. Anadolu Bil Meslek Yüksekokulu Dergisi, vol. 21, no. 73, July 2026, pp. 171-95, https://izlik.org/JA86UN57UM.
Vancouver
1.Neslihan Gündoğan, Buket İşler. MFCCFusionnet: A CNN–DNN feature-fusion model for speech emotion recognition using MFCC representations. ABMYO Dergisi [Internet]. 2026 Jul. 1;21(73):171-95. Available from: https://izlik.org/JA86UN57UM



All site content, except where otherwise noted, is licensed under a Creative Common Attribution Licence. (CC-BY-NC 4.0)