Zero Crossing Rate (ZCR) Feature Extraction in Speech Emotion Recognition Using a Convolutional Neural Network

Authors

  • Silviana Widya Lestari Institut Teknologi dan Bisnis Widya Gama Lumajang
  • Maysas Yafi Urrochman Institut Teknologi dan Bisnis Widya Gama Lumajang
  • Cahyasari Kartika Murni Institut Teknologi dan Bisnis Widya Gama Lumajang

DOI:

https://doi.org/10.30741/jid.v5i1.2083

Keywords:

Speech Emotion Recognition, Zero Crossing Rate, Convolutional Neural Network, Feature Extraction, Deep Learning

Abstract

Speech Emotion Recognition (SER) plays a pivotal role in enhancing human-computer interaction (HCI), healthcare, and automated feedback systems (Hashem et al., 2023). However, extracting relevant acoustic features from complex speech signals remains a significant challenge (Lieskovská et al., 2021). This paper investigates the effectiveness of the Zero Crossing Rate (ZCR) feature extraction method combined with a 1D Convolutional Neural Network (CNN) architecture for speech emotion classification. By evaluating emotional speech data across benchmark datasets (RAVDESS, CREMA, SAVEE, and TESS), audio signals were preprocessed using data augmentation techniques, including noise injection, time stretching, pitch shifting, and time shifting. The ZCR feature representation was fed into a CNN model to classify seven core emotional states: anger, disgust, fear, happiness, sadness, surprise, and neutral. The proposed ZCR-CNN approach achieved a classification accuracy of 61.96% with a training loss of 0.8214 and testing loss of 0.8123. Although ZCR effectively captures high-frequency and noisiness variations indicative of specific emotional states, its performance highlights the trade-offs of relying on single-domain time features compared to spectral features (Singh & Goel, 2022).

References

Alluhaidan, F., et al. (2023). Speech emotion recognition using machine learning and deep learning techniques. Journal of Voice, 37(4), 512-525.

Anvarjon, T., Mustaqeem, & Kwon, S. (2020). Deep-SER: Real-time speech emotion recognition using deep convolutional neural networks. IEEE Access, 8, 11472-11484.

Hashem, M., et al. (2023). Speech emotion recognition approaches: A systematic literature review. Applied Sciences, 13(2), 890.

Issa, D., Demirci, M. F., & Yazici, A. (2020). Speech emotion recognition with deep convolutional neural networks. Biomedical Signal Processing and Control, 59, 101894.

Lieskovská, E., et al. (2021). Review on feature extraction and classification algorithms for speech emotion recognition. Sensors, 21(6), 2134.

Mustaqeem, & Kwon, S. (2020). A CNN-assisted lightweight framework for speech emotion recognition. IEEE Access, 8, 115450-115462.

Pandey, S. K., Shekhawat, H. S., & Prasanna, S. R. M. (2019). Deep learning techniques for speech emotion recognition: A review. IEEE Transactions on Cognitive and Developmental Systems, 11(4), 450-463.

Singh, P., & Goel, S. (2022). Feature extraction techniques for speech emotion recognition: A comparative study. Computer Science Review, 44, 100465.

Downloads

Published

2026-10-01

How to Cite

Lestari, S. W., Urrochman, M. Y., & Murni, C. K. (2026). Zero Crossing Rate (ZCR) Feature Extraction in Speech Emotion Recognition Using a Convolutional Neural Network. Journal of Informatics Development, 5(1), 41–45. https://doi.org/10.30741/jid.v5i1.2083

Most read articles by the same author(s)

<< < 1 2