Emotion Recognition From Speech With Recurrent Neural Networks
arXiv:1701.08071
Abstract
In this paper the task of emotion recognition from speech is considered. Proposed approach uses deep recurrent neural network trained on a sequence of acoustic features calculated over small speech intervals. At the same time special probabilistic-nature CTC loss function allows to consider long utterances containing both emotional and neutral parts. The effectiveness of such an approach is shown in two ways. Firstly, the comparison with recent advances in this field is carried out. Secondly, human performance on the same task is measured. Both criteria show the high quality of the proposed method.
References in corpus (4)
Cited by in corpus (10)
- Multi-Modal Emotion recognition on IEMOCAP Dataset using Deep Learning
- Deep Learning based Emotion Recognition System Using Speech Features and Transcriptions
- Learning Speech Emotion Representations in the Quaternion Domain
- Reusing Neural Speech Representations for Auditory Emotion Recognition
- Designing and Evaluating Speech Emotion Recognition Systems: A reality check case study with IEMOCAP
- A Transfer Learning Method for Speech Emotion Recognition from Automatic Speech Recognition
- Emotions Don't Lie: An Audio-Visual Deepfake Detection Method Using Affective Cues
- Learning Discriminative features using Center Loss and Reconstruction as Regularizer for Speech Emotion Recognition
- Multi-Window Data Augmentation Approach for Speech Emotion Recognition
- Multi-Modal Emotion Detection with Transfer Learning