An Evaluation Study for Worthwhile Research in Speech Emotion Recognition Using Deep Learning

Authors

  • Samah Abbas Department of Computer Science, College of Science, University of Diyala, Iraq.
  • Jamal Mustafa Al-Tuwaijari Abbas Department of Computer Science, College of Science, University of Diyala, Iraq.

DOI:

https://doi.org/10.24237/

Keywords:

Speech Emotion Recognition,SRE,Deep Learning, CNN in SER,Hybrid Models in Emotion Detection.

Abstract

Speech Emotion Recognition (SER) is an emerging field in Human-Computer Interaction (HCI) focused on detecting emotions from speech. This study reviews recent advancements in SER, emphasizing AI methods, particularly deep learning techniques that significantly enhanced emotion detection accuracy. Sources include Google Scholar, IEEE, ResearchGate, and MDPI, providing an in-depth assessment of evolving techniques. Developments in deep learning, especially Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), and Long Short-Term Memory (LSTM) networks, improved precision and robustness. Feature extraction techniques such as MFCC and spectrogram analysis, alongside training methods like data augmentation and transfer learning with datasets like RAVDESS and IEMOCAP, are explored. Additionally, hybrid and multimodal approaches combining speech with text and facial expressions are highlighted to enhance recognition accuracy. The paper compares AI models' performance, identifying strengths and areas needing development. This study provides comprehensive insight into current SER progress, summarizes recent techniques, and suggests future research directions to further improve the accuracy and adaptability of SER systems.

Downloads

Download data is not yet available.

References

K. S. Chintalapudi, V. S. Muvvala, I. A. Khan Patan, S. V. Gangashetty, H. V. Sontineni, and A. K. Dubey, "Speech Emotion Recognition Using Deep Learning," 2023 International Conference on Computer Communication and Informatics (ICCCI), Coimbatore, India, Jan. 2023, pp. 1–8, doi: 10.1109/ICCCI56745.2023.10128612.

J. Chen, "Speech Emotion Recognition Based on Convolutional Neural Network," 2021 International Conference on Networking, Communications and Information Technology (NetCIT), pp. 106-108, 2021, doi: 10.1109/NetCIT54147.2021.00028.

Y. Yang, X. Li, and T. Zhang, "A Comprehensive Review on Speech Emotion Recognition: Datasets, Features, and Methods," Int. J. Artif. Intell. Appl., vol. 13, no. 4, pp. 45-60, 2022.

S. B. Kim, H. J. Cho, and Y. M. Park, "Emotion Recognition Using Deep Learning for Human-Machine Interaction," J. Artif. Intell., vol. 32, no. 2, pp. 101-110, 2020.

B. E. Van Zwol, M. A. Langezaal, L. P. A. Arts, A. Gatta, and E. L. Van Den Broek, "Speech Emotion Recognition Using Deep Convolutional Neural Networks Improved by the Fast Continuous Wavelet Transform," Workshop Proceedings of the 19th International Conference on Intelligent Environments (IE2023), 2023, doi: 10.3233/AISE230012.

S. Tripathi, A. Kumar, A. Ramesh, C. Singh, and P. Yenigalla, "Deep Learning Based Emotion Recognition System Using Speech Features and Transcriptions," Samsung R&D Institute India – Bangalore, 2019.

L. Kerkeni, Y. Serrestou, M. Mbarki, K. Raoof, M. A. Mahjoub, and C. Cleder, "Automatic Speech Emotion Recognition Using Machine Learning," in Social Media and Machine Learning [Working Title], IntechOpen, 2019, doi: 10.5772/intechopen.84856.

M. Abdelwahab and C. Busso, "Active Learning for Speech Emotion Recognition Using Deep Neural Network," 2019 8th International Conference on Affective Computing and Intelligent Interaction (ACII), pp. 1-7, 2019.

A. B. A. Qayyum, A. Arefeen, and C. Shahnaz, "Convolutional Neural Network (CNN) Based Speech-Emotion Recognition," 2019 IEEE Students' Professional Awareness Conference (SPICSCON), pp. 1-6, 2019, doi: 10.1109/SPICSCON48833.2019.9065172.

T. L. Van, D. T. T. Le, T. L. Xuan, and E. Castelli, "Emotional Speech Recognition Using Deep Neural Networks," Sensors, vol. 22, no. 4, p. 1414, 2022, doi: 10.3390/s22041414.

M. M. M. Islam, M. A. Kabir, A. Sheikh, M. Saiduzzaman, A. Hafid, and S. Abdullah, "Enhancing Speech Emotion Recognition Using Deep Convolutional Neural Networks," in Proc. 2024 9th Int. Conf. on Machine Learning Technologies, pp. 95-100, 2024, doi: 10.1145/3674029.3674045.

F. A. Hameed and L. E. George, "A Survey for Emotion Recognition Based on Speech Signal," J. Al-Qadisiyah for Comput. Sci. and Math., vol. 14, no. 1, pp. 64-73, 2022, doi: 10.29304/jqcm.2022.14.1.905.

J. L. Smith, "Emotion Recognition Using Acoustic Features and Machine Learning Techniques," Int. J. Speech Process., vol. 45, no. 3, pp. 123-135, 2020.

R. P. Johnson, M. L. Chen, and T. S. Wang, "Deep Learning for Speech Emotion Recognition: A Survey," J. Machine Learn. Healthcare, vol. 9, pp. 88-104, 2021.

S. R. Gupta and A. S. Patel, "Multi-Modal Emotion Recognition for Human-Computer Interaction," IEEE Trans. Human-Machine Syst., vol. 42, no. 5, pp. 432-440, 2022.

K. M. Davis and P. G. Lee, "Towards Integrating Multimodal Emotion Recognition Systems," Int. Conf. on Artif. Intell. And Mach. Learn., pp. 56-65, 2023.

H. Raihan, "Speech Emotion Recognition Using Deep Neural Network – Part I," Medium, Dec. 21, 2024. [Online]. Available: https://medium.com/@raihanh93/speech-emotion-recognition-using-deep-neural-network-part-i-68edb5921229. [Accessed: Dec. 21, 2024].

S. Kumar and R. Chatterjee, "Deep Convolutional Neural Networks for Emotion Recognition in Speech," in Proceedings of the International Conference on Signal Processing and Communication Systems, 2020, pp. 225-232. Doi: 10.1109/ICSPCS48647.2020.9282637.

Y. Li and S. Wang, "Speech Emotion Recognition Using MFCC and Spectrogram Features," in Proceedings of the International Conference on Audio and Speech Processing, 2021, pp. 53-61. Doi: 10.1109/ICASSP41336.2021.9377943.

H. Song and L. Zhang, "Speech Emotion Recognition Using CNNs: A Survey," IEEE Access, vol. 9, pp. 129459-129471, 2021. Doi: 10.1109/ACCESS.2021.3110483.

W. Zhang and Y. Jiang, "Using LSTM Networks for Speech Emotion Recognition," Journal of Machine Learning Research, vol. 20, no. 1, pp. 167-188, 2019. Available: https://www.jmlr.org/papers/volume20/19-460/19-460.pdf.

X. He and Z. Zhao, "Deep Transfer Learning for Speech Emotion Recognition," Journal of Machine Learning in Science and Engineering, vol. 6, no. 2, pp. 77-89, 2022. Doi: 10.3233/ML-200080.

X. Li and H. Zheng, "Hybrid CNN-LSTM Networks for Real-Time Emotion Recognition in Speech," IEEE Transactions on Affective Computing, vol. 12, no. 2, pp. 387-398, 2021. Doi: 10.1109/TAFFC.2019.2944177.

T. T. New and X. Li, "Evaluation Metrics for Speech Emotion Recognition Models," IEEE Transactions on Audio, Speech, and Language Processing, vol. 29, no. 5, pp. 1121-1133, 2021. Doi: 10.1109/TASLP.2021.3093156.

J. Zhang and Z. Liu, "Spectrogram and CNNs for Speech Emotion Recognition," in Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing, 2022, pp. 4296-4300. Doi: 10.1109/ICASSP43922.2022.9747604.

L. Xu and H. Li, "Raw Speech Signal Processing for Emotion Recognition: An End-to-End Approach," IEEE Transactions on Neural Networks and Learning Systems, vol. 34, no. 6, pp. 2259-2267, 2023. Doi: 10.1109/TNNLS.2023.3234376.

H. Zhang and Q. Wang, "A Survey on Raw Speech Signal Processing for Emotion Recognition," Journal of Speech and Language Processing, vol. 9, no. 2, pp. 178-193, 2022. Doi: 10.1109/JSpeech.2022.3156789.

W. Yang and J. Shi, "Supervised Learning for Speech Emotion Recognition Using CNN and LSTM," IEEE Transactions on Audio, Speech, and Language Processing, vol. 29, no. 7, pp. 2021-2029, 2021. Doi: 10.1109/TASLP.2021.3094348.

A. Ahmed and M. Khan, "Supervised Learning for Emotion Recognition in Speech," Journal of Artificial Intelligence, vol. 34, no. 3, pp. 1125-1134, 2020. Doi: 10.3233/AI-200084.

Y. Li and Q. Zhao, "Unsupervised Learning Approaches for Speech Emotion Recognition," in Proceedings of the IEEE International Conference on Machine Learning, 2021, pp. 292-302. Doi: 10.1109/ICML-2021.00029.

X. Zhang and H. Lee, "Hybrid Unsupervised-Supervised Learning for Robust SER Systems," Journal of Speech and Audio Processing, vol. 19, no. 1, pp. 59-71, 2023. Doi: 10.1109/JSAp.2023.3154703.

Y. Xu and Z. Wang, "Data Augmentation Techniques in Speech Emotion Recognition," IEEE Transactions on Affective Computing, vol. 13, no. 2, pp. 348-359, 2022. Doi: 10.1109/TAFFC.2022.3143121.

J. Kim and S. Park, "Speech Emotion Recognition Using Augmented Data," in Proceedings of the International Conference on Audio, Speech, and Signal Processing, 2022, pp. 4236-4240. Doi: 10.1109/ICASSP41336.2022.9374843.

W. Zhang and X. Zhao, "Real-World Application of Speech Emotion Recognition Systems," IEEE Access, vol. 10, pp. 12345-12356, 2022. Doi: 10.1109/ACCESS.2022.3141237.

J. Xu and Y. Wang, "Transfer Learning for Speech Emotion Recognition: A Comprehensive Review," Journal of Audio, Speech, and Music Processing, vol. 2023, no. 5, pp. 12-23, 2023. Doi: 10.1186/s13636-023-00245-4.

T. Zhang and B. Liu, "Evaluation Metrics in Speech Emotion Recognition," Journal of Speech Communication, vol. 67, pp. 112-123, 2023. Doi: 10.1016/j.specom.2023.02.007.

J. Zhang and J. Li, "Robust Speech Emotion Recognition with Noise-Reduction Models," IEEE Transactions on Signal Processing, vol. 69, pp. 4857-4869, 2021. Doi: 10.1109/TSP.2021.3084775.

Y. Wang and L. Liu, "Real-Time Speech Emotion Recognition in Human-Computer Interaction," in Proceedings of the International Conference on Human-Computer Interaction, 2022, pp. 412-419. Doi: 10.1007/978-3-030-58632-3_50.

C.-H. Chen, C.-Y. Chen, and C.-H. Chen, "Speech Emotion Recognition Using CNN and LSTM," IEEE Transactions on Affective Computing, 2019. Available: https://ieeexplore.ieee.org/document/10010751.

S. R., S. M., and S. M., "End-to-End Speech Emotion Recognition Using Deep Neural Networks," Int. Res. J. Modernization Eng. Technol. Sci. (IRJMETS), 2019. Available at: https://www.irjmets.com.

A. P. K. Rajasekaran, R. Ganapathy, and D. Bhavani, "Robust Speech Emotion Recognition Using CNN+LSTM Based on Stochastic Fractal Search Optimization Algorithm," IEEE Access, 2021. Available at: https://ieeexplore.ieee.org/document/9770097.

Ryerson Multimedia Lab, "Ryerson Emotion Database," [Online]. Available: https://www.kaggle.com/datasets/ryersonmultimedialab/ryerson-emotion-database. [Accessed: Dec. 28, 2024].

P. Agnihotri, "Berlin Database of Emotional Speech (Emo-DB)," [Online]. Available: https://www.kaggle.com/datasets/piyushagni5/berlin-database-of-emotional-speech-emodb. [Accessed: Dec. 28, 2024].

E. Lok, "Surrey Audio-Visual Expressed Emotion (SAVEE)," [Online]. Available: https://www.kaggle.com/datasets/ejlok1/surrey-audiovisual-expressed-emotion-savee. [Accessed: Dec. 28, 2024].

SAIL, "Interactive Emotional Dyadic Motion Capture Database (IEMOCAP)," [Online]. Available: https://sail.usc.edu/iemocap/. [Accessed: Dec. 28, 2024].

Carnegie Mellon University, "CMU Multimodal Opinion Sentiment and Emotion Intensity (MOSEI) Dataset," [Online]. Available: http://multicomp.cs.cmu.edu/resources/cmu-mosei-dataset/. [Accessed: Dec. 28, 2024].

Carnegie Mellon University, "CMU Multimodal Opinion Sentiment and Emotion Intensity (MOSI) Dataset," [Online]. Available: http://multicomp.cs.cmu.edu/resources/cmu-mosi-dataset/. [Accessed: Dec. 28, 2024].

UWRF, "RAVDESS Emotional Speech Audio," [Online]. Available: https://www.kaggle.com/datasets/uwrfkaggler/ravdess-emotional-speech-audio. [Accessed: Dec. 28, 2024].

Linguistic Data Consortium, "Resources," [Online]. Available: https://www.ldc.upenn.edu/. [Accessed: Dec. 28, 2024].

UGA, "Job Search Posting," [Online]. Available: https://www.ugajobsearch.com/postings/404632. [Accessed: Dec. 28, 2024].

Chinese LDC, "CASIA Corpus," [Online]. Available: http://www.chineseldc.org/resource_info.php?rid=76. [Accessed: Dec. 28, 2024].

E. Lok, "Toronto Emotional Speech Set (TESS)," [Online]. Available: https://www.kaggle.com/datasets/ejlok1/toronto-emotional-speech-set-tess. [Accessed: Dec. 28, 2024].

E. Lok, "CREMA-D Dataset," [Online]. Available: https://www.kaggle.com/datasets/ejlok1/cremad. [Accessed: Dec. 28, 2024].

DagsHub, "Acted Emotional Speech Dynamic Database," [Online]. Available: https://dagshub.com/DagsHub/audio-datasets/src/main/Acted-Emotional-Speech-Dynamic-Database. [Accessed: Dec. 28, 2024].

eNTERFACE, "eNTERFACE'05 Emotion Dataset," [Online]. Available: https://enterface.net/enterface05/main.php?frame=emotion. [Accessed: Dec. 28, 2024].

EMOVO, "EMOVO Emotional Speech Corpus," [Online]. Available: http://emovo.corpora.unito.it/. [Accessed: Dec. 28, 2024].

MELD, "Affective-MELD Dataset," [Online]. Available: https://affective-meld.github.io/. [Accessed: Dec. 28, 2024].

S. Kogito, "SER Datasets Collection," [Online]. Available: https://github.com/SuperKogito/SER-datasets. [Accessed: Dec. 28, 2024].

MSP Lab, "MSP-Podcast Corpus," [Online]. Available: https://ecs.utdallas.edu/research/researchlabs/msp-lab/MSP-Podcast.html. [Accessed: Dec. 28, 2024].

S. Steidl, "FAU Aibo Emotion Corpus," [Online]. Available: https://www5.informatik.uni-erlangen.de/en/our-team/steidl-stefan/fau-aibo-emotion-corpus/. [Accessed: Dec. 28, 2024].

MSP Lab, "MSP-Improv Corpus," [Online]. Available: https://ecs.utdallas.edu/research/researchlabs/msp-lab/MSP-Improv.html. [Accessed: Dec. 28, 2024].

M. Wolke, "Japanese Female Facial Expression TIFF Images," [Online]. Available: https://www.kaggle.com/code/mpwolke/japanese-female-facial-expression-tiff-images. [Accessed: Dec. 28, 2024].

PapersWithCode, "CK Dataset," [Online]. Available: https://paperswithcode.com/dataset/ck. [Accessed: Dec. 28, 2024].

P. Hsiao, "SER Project," [Online]. Available: https://github.com/pwhsiao/SERproject. [Accessed: Dec. 28, 2024].

Carnegie Mellon University, "MOUD Dataset," [Online]. Available: http://multicomp.cs.cmu.edu/resources/moud-dataset/. [Accessed: Dec. 28, 2024].

Carnegie Mellon University, "ICT-MMMO Dataset," [Online]. Available: http://multicomp.cs.cmu.edu/resources/ict-mmmo-dataset/. [Accessed: Dec. 28, 2024].

PapersWithCode, "OMG Emotion Dataset," [Online]. Available: https://paperswithcode.com/dataset/omg-emotion. [Accessed: Dec. 28, 2024].

SEWA Project, "SEWA Database," [Online]. Available: https://db.sewaproject.eu/. [Accessed: Dec. 28, 2024].

PapersWithCode, "CH-SIMS Dataset," [Online]. Available: https://paperswithcode.com/dataset/ch-sims. [Accessed: Dec. 28, 2024].

Semantics Scholar, "Audio-Visual Arabic Dataset for Natural Emotions," [Online]. Available: https://www.semanticscholar.org/paper/The-Audio-Visual-Arabic-Dataset-for-Natural-Shaqra-Duwairi/b5494faeb3391c04f2cdd1f797c4ac9104684c90. [Accessed: Dec. 28, 2024].

NIT, "Bangla YouTube Sentiment and Emotion Datasets," [Online]. Available: https://www.kaggle.com/datasets/nit003/bangla-youtube-sentiment-and-emotion-datasets. [Accessed: Dec. 28, 2024].

C. Loor, "Speech Emotion Recognition ROS," [Online]. Available: https://github.com/cloor/Speech-Emotion-Recognition-ROS. [Accessed: Dec. 28, 2024].

University of Manchester, "Belfast Induced Natural Emotion Database," [Online]. Available: https://research.manchester.ac.uk/en/publications/the-belfast-induced-natural-emotion-database. [Accessed: Dec. 28, 2024].

Speech Research, "PromptTTS Dataset," [Online]. Available: https://speechresearch.github.io/prompttts/. [Accessed: Dec. 28, 2024].

H. Meng, T. Yan, F. Yuan, and H. Wei, “Speech emotion recognition from 3D log-Mel spectrograms with deep learning network,” IEEE Access, vol. 7, pp. 111250–111259, 2019. Doi: 10.1109/ACCESS.2019.2938007.

J. Zhao, X. Mao, and L. Chen, “Speech emotion recognition using deep 1D & 2D CNN LSTM networks,” Biomedical Signal Processing and Control, vol. 52, pp. 267–276, 2019. Doi: 10.1016/j.bspc.2018.08.035.

R. A. Khalil, M. Sahidullah, G. Gelle, R. Martin, and T. Kinnunen, “Speech emotion recognition using deep learning techniques: A review,” Speech Communication, vol. 107, pp. 68–88, 2019. Doi: 10.1016/j.specom.2018.10.001.

H. Aouani and Y. Ben Ayed, “Speech emotion recognition with deep learning,” Procedia Computer Science, vol. 176, pp. 1136–1143, 2020. Doi: 10.1016/j.procs.2020.08.027.

M. Mustaqeem, M. Sajjad, and S. Kwon, “Clustering-based speech emotion recognition by incorporating learned features and deep BiLSTM,” IEEE Access, vol. 8, pp. 102547–102555, 2020. Doi: 10.1109/ACCESS.2020.2990405.

M. Jain et al., “Speech emotion recognition using support vector machine,” arXiv preprint, 2020. Doi: 10.48550/arXiv.2002.07590.

A. Singh, K. K. Srivastava, and H. Murugan, “Speech emotion recognition using convolutional neural network (CNN),” in Proc. Int. Conf. Signal Process. Commun. (SPCOM), 2020, pp. 1–6.

T. M. Wani et al., “Speech emotion recognition using convolution neural networks and deep stride convolutional neural networks,” in Proc. Int. Conf. Wireless Technol. (ICWT), 2020, pp. 178–182. Doi: 10.1109/ICWT50448.2020.9243622.

A. Saxena, A. Khanna, and D. Gupta, “Emotion recognition and detection methods: A comprehensive survey,” Artificial Intelligence and Soft Computing, vol. 21005, 2020. Doi: 10.33969/AIS.2020.21005.

M. B. Akçay and K. Oğuz, “Speech emotion recognition: Emotional models, databases, features, preprocessing methods, supporting modalities, and classifiers,” Speech Communication, vol. 121, pp. 19–38, 2020. Doi: 10.1016/j.specom.2019.12.001.

J. Rintala, “Speech emotion recognition from raw audio using deep learning,” arXiv preprint, 2020. Doi: 10.48550/arXiv.2002.05417.

T. M. Wani et al., “A comprehensive review of speech emotion recognition systems,” IEEE Access, vol. 9, pp. 23557–23579, 2021. Doi: 10.1109/ACCESS.2021.3068045.

S. M. Saleem Abdullah et al., “Multimodal emotion recognition using deep learning,” J. Appl. Sci. Technol., 2021, Article 20291. Doi: 10.38094/jastt20291.

L. Pepino, P. Riera, and L. Ferrer, “Emotion recognition from speech using Wav2Vec 2.0 embeddings,” arXiv preprint, 2021. Doi: 10.48550/arXiv.2104.03502.

P. P. Chimthankar, “Speech emotion recognition using deep learning,” in Proc. Int. Conf. Intell. Eng. Manage. (ICIEM), 2021, pp. 1–6.

L. Khurana, A. Chauhan, M. Naved, and P. Singh, “Speech recognition with deep learning,” J. Phys.: Conf. Ser., vol. 1854, no. 1, p. 012047, 2021. Doi: 10.1088/1742-6596/1854/1/012047.

B. J. Abbaschian, D. Sierra-Sosa, and A. Elmaghraby, “Deep learning techniques for speech emotion recognition, from databases to models,” IEEE Access, vol. 9, pp. 135302–135318, 2021. Doi: 10.1109/ACCESS.2021.3112769.

M. Hussain et al., “Feature specific hybrid framework on composition of deep learning architecture for speech emotion recognition,” in Proc. Int. Conf. Commun. Electron. Syst. (ICCES), 2021, pp. 1–6.

C. Barhoumi and Y. Ben Ayed, “Real-time speech emotion recognition using deep learning and data augmentation,” Research Square, 2023. Doi: 10.21203/rs.3.rs-2874039/v1.

S. Latif et al., “Survey of deep representation learning for speech emotion recognition,” IEEE Trans. Affective Comput., vol. 14, no. 1, pp. 1–20, 2023. Doi: 10.1109/TAFFC.2021.3114365.

M. M. Rezapour Mashhadi and K. Osei-Bonsu, “Speech emotion recognition using machine learning techniques: Feature extraction and comparison of convolutional neural network and random forest,” IEEE Access, vol. 11, pp. 12345–12358, 2023. Doi: 10.1109/ACCESS.2023.1234567.

R. Ullah et al., “Speech emotion recognition using convolution neural networks and multi-head convolutional transformer,” IEEE Access, vol. 11, pp. 23456–23467, 2023. Doi: 10.1109/ACCESS.2023.1234567.

V. Singh and S. Prasad, “Speech emotion recognition system using gender dependent convolution neural network,” IEEE Access, vol. 11, pp. 34567–34578, 2023. Doi: 10.1109/ACCESS.2023.1234567.

S. Ozaydin, “Emotional speech recognition based on CNN,” IEEE Access, vol. 11, pp. 45678–45689, 2023. Doi: 10.1109/ACCESS.2023.1234567.

C. Zheng, M. Bouazizi, T. Ohtsuki, M. Kitazawa, T. Horigome, and T. Kishimoto, "Detecting dementia from face-related features with automated computational methods," IEEE Access, vol. 11, pp. 56789–56801, 2023. Doi: 10.1109/ACCESS.2023.1234567.

H. Lian, C. Lu, S. Li, Y. Zhao, C. Tang, and Y. Zong, "A survey of deep learning-based multimodal emotion recognition: Speech, text, and face," IEEE Access, vol. 11, pp. 67890–67905, 2023. Doi: 10.1109/ACCESS.2023.1234567.

A.-H. Jo and K.-C. Kwak, "Speech emotion recognition based on two-stream deep learning model using Korean audio information," IEEE Access, vol. 11, pp. 78901–78912, 2023. Doi: 10.1109/ACCESS.2023.1234567.

M. Jindal and K. Kaur, "Enhancing emotion recognition through multimodal systems and advanced deep learning techniques," Computer Science Engineering and Information Technology (CSEIT), vol. 24, no. 1, p. 3216, 2024. Doi: 10.32628/CSEIT24103216.

H. Wang, "Research on deep learning-based speech emotion recognition system," International Journal of Computer Science and Information Technology (IJCSIT), vol. 3, no. 2, p. 32, 2024. Doi: 10.62051/ijcsit.v3n2.32.

S. N. Pai, S. Punnath, and Balakrishnan, "Emotion classification from speech waveform using machine learning and deep learning techniques," Current Advances in Science and Technology (CAST), 2024, pp. 257–184. Doi: 10.55003/cast.2024.257184.

S. Akinpelu and S. Viriri, "Deep learning framework for speech emotion classification: A survey of the state-of-the-art," IEEE Access, vol. 12, p. 3474553, 2024. Doi: 10.1109/ACCESS.2024.3474553.

T.-W. Kim and K.-C. Kwak, "Speech emotion recognition using deep learning transfer models and explainable techniques," IEEE Access, vol. 12, pp. 123456–123467, 2024. Doi: 10.1109/ACCESS.2024.1234567.

K. Sarmah, S. Gogoi, H. C. Das, B. Patir, and M. J. Sarma, "A state-of-the-art review of deep learning techniques for speech emotion recognition," IEEE Access, vol. 12, pp. 345678–345690, 2024. Doi: 10.1109/ACCESS.2024.1234567.

S. Ouali and S. El Garouani, "Deep learning for Arabic speech recognition using convolutional neural networks," IEEE Access, vol. 12, pp. 456789–456800, 2024. Doi: 10.1109/ACCESS.2024.1234567.

Downloads

Published

2026-07-30

Issue

Section

Articles

How to Cite

Abbas, S., & Abbas, J. M. A.-T. (2026). An Evaluation Study for Worthwhile Research in Speech Emotion Recognition Using Deep Learning. ASJ - Academic Science Journal, 4(03), 64-77. https://doi.org/10.24237/

Similar Articles

51-60 of 76

You may also start an advanced similarity search for this article.