Research Article

END-TO-END AUTOMATIC MUSIC TRANSCRIPTION OF POLYPHONIC QANUN AND OUD MUSIC USING DEEP NEURAL NETWORK

Volume: 25 Number: 3 September 30, 2024
EN

END-TO-END AUTOMATIC MUSIC TRANSCRIPTION OF POLYPHONIC QANUN AND OUD MUSIC USING DEEP NEURAL NETWORK

Abstract

This paper introduces an automatic music transcription model using Deep Neural Networks (DNNs), focusing on simulating the "trained ear" in music. It advances the field of signal processing and music technology, particularly in multi-instrument transcription involving traditional Turkish instruments, Qanun and Oud. Those instruments have unique timbral characteristics with early decay periods. The study involves generating basic combinations of multi-pitch datasets, training the DNN model on this data, and demonstrating its effectiveness in transcribing two-part compositions with high accuracy and F1 measures. The model's training involves understanding the fundamental characteristics of individual instruments, enabling it to identify and isolate complex patterns in mixed compositions. The primary goal is to empower the model to distinguish and analyze individual musical components, thereby enhancing applications in music production, audio engineering, and education

Keywords

Music transcription, Signal processing, Deep learning, Constant Q transform

References

  1. [1] Benetos E, Dixon S, Duan Z, Ewert S. Automatic Music Transcription: An Overview. IEEE Signal Process Mag. 2019;36(1):20-30. doi:10.1109/MSP.2018.2869928
  2. [2] Bertin N, Badeau R, Richard G. Blind Signal Decompositions for Automatic Transcription of Polyphonic Music: NMF and K-SVD on the Benchmark. In: 2007 IEEE International Conference on Acoustics, Speech and Signal Processing - ICASSP ’07. IEEE; 2007:I-65-I-68. doi:10.1109/ICASSP.2007.366617
  3. [3] Ansari S, Alatrany AS, Alnajjar KA, et al. A survey of artificial intelligence approaches in blind source separation. Neurocomputing. 2023;561:126895. doi:https://doi.org/10.1016/j.neucom.2023.126895
  4. [4] Munoz-Montoro AJ, Carabias-Orti JJ, Cabanas-Molero P, Canadas-Quesada FJ, Ruiz-Reyes N. Multichannel Blind Music Source Separation Using Directivity-Aware MNMF With Harmonicity Constraints. IEEE Access. 2022;10:17781-17795. doi:10.1109/ACCESS.2022.3150248
  5. [5] Uhlich S, Giron F, Mitsufuji Y. Deep neural network based instrument extraction from music. In: ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings. ; 2015. doi:10.1109/ICASSP.2015.7178348
  6. [6] Nishikimi R, Nakamura E, Goto M, Yoshii K. Audio-to-score singing transcription based on a CRNN-HSMM hybrid model. APSIPA Trans Signal Inf Process. 2021;10(1):e7. doi:10.1017/ATSIP.2021.4
  7. [7] Sigtia S, Benetos E, Boulanger-Lewandowski N, Weyde T, d’Avila Garcez AS, Dixon S. A hybrid recurrent neural network for music transcription. In: 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP). IEEE; 2015:2061-2065. doi:10.1109/ICASSP.2015.7178333
  8. [8] Sigtia S, Benetos E, DIxon S. An end-to-end neural network for polyphonic piano music transcription. IEEE/ACM Trans Audio Speech Lang Process. 2016;24(5):927-939. doi:10.1109/TASLP.2016.2533858
  9. [9] Seetharaman P, Pishdadian F, Pardo B. Music/Voice separation using the 2D fourier transform. In: IEEE Workshop on Applications of Signal Processing to Audio and Acoustics. ; 2017. doi:10.1109/WASPAA.2017.8169990
  10. [10] Huang P Sen, Chen SD, Smaragdis P, Hasegawa-Johnson M. Singing-voice separation from monaural recordings using robust principal component analysis. In: ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings. ; 2012. doi:10.1109/ICASSP.2012.6287816
APA
Germen, E., & Karadoğan, C. (2024). END-TO-END AUTOMATIC MUSIC TRANSCRIPTION OF POLYPHONIC QANUN AND OUD MUSIC USING DEEP NEURAL NETWORK. Eskişehir Technical University Journal of Science and Technology A - Applied Sciences and Engineering, 25(3), 442-455. https://doi.org/10.18038/estubtda.1467350
AMA
1.Germen E, Karadoğan C. END-TO-END AUTOMATIC MUSIC TRANSCRIPTION OF POLYPHONIC QANUN AND OUD MUSIC USING DEEP NEURAL NETWORK. Estuscience - Se. 2024;25(3):442-455. doi:10.18038/estubtda.1467350
Chicago
Germen, Emin, and Can Karadoğan. 2024. “END-TO-END AUTOMATIC MUSIC TRANSCRIPTION OF POLYPHONIC QANUN AND OUD MUSIC USING DEEP NEURAL NETWORK”. Eskişehir Technical University Journal of Science and Technology A - Applied Sciences and Engineering 25 (3): 442-55. https://doi.org/10.18038/estubtda.1467350.
EndNote
Germen E, Karadoğan C (September 1, 2024) END-TO-END AUTOMATIC MUSIC TRANSCRIPTION OF POLYPHONIC QANUN AND OUD MUSIC USING DEEP NEURAL NETWORK. Eskişehir Technical University Journal of Science and Technology A - Applied Sciences and Engineering 25 3 442–455.
IEEE
[1]E. Germen and C. Karadoğan, “END-TO-END AUTOMATIC MUSIC TRANSCRIPTION OF POLYPHONIC QANUN AND OUD MUSIC USING DEEP NEURAL NETWORK”, Estuscience - Se, vol. 25, no. 3, pp. 442–455, Sept. 2024, doi: 10.18038/estubtda.1467350.
ISNAD
Germen, Emin - Karadoğan, Can. “END-TO-END AUTOMATIC MUSIC TRANSCRIPTION OF POLYPHONIC QANUN AND OUD MUSIC USING DEEP NEURAL NETWORK”. Eskişehir Technical University Journal of Science and Technology A - Applied Sciences and Engineering 25/3 (September 1, 2024): 442-455. https://doi.org/10.18038/estubtda.1467350.
JAMA
1.Germen E, Karadoğan C. END-TO-END AUTOMATIC MUSIC TRANSCRIPTION OF POLYPHONIC QANUN AND OUD MUSIC USING DEEP NEURAL NETWORK. Estuscience - Se. 2024;25:442–455.
MLA
Germen, Emin, and Can Karadoğan. “END-TO-END AUTOMATIC MUSIC TRANSCRIPTION OF POLYPHONIC QANUN AND OUD MUSIC USING DEEP NEURAL NETWORK”. Eskişehir Technical University Journal of Science and Technology A - Applied Sciences and Engineering, vol. 25, no. 3, Sept. 2024, pp. 442-55, doi:10.18038/estubtda.1467350.
Vancouver
1.Emin Germen, Can Karadoğan. END-TO-END AUTOMATIC MUSIC TRANSCRIPTION OF POLYPHONIC QANUN AND OUD MUSIC USING DEEP NEURAL NETWORK. Estuscience - Se. 2024 Sep. 1;25(3):442-55. doi:10.18038/estubtda.1467350