End-to-End Multimodal Emotion Recognition using Deep Neural Networks
arXiv:1704.08619 · doi:10.1109/JSTSP.2017.2764438
Abstract
Automatic affect recognition is a challenging task due to the various modalities emotions can be expressed with. Applications can be found in many domains including multimedia retrieval and human computer interaction. In recent years, deep neural networks have been used with great success in determining emotional states. Inspired by this success, we propose an emotion recognition system using auditory and visual modalities. To capture the emotional content for various styles of speaking, robust features need to be extracted. To this purpose, we utilize a Convolutional Neural Network (CNN) to extract features from the speech, while for the visual modality a deep residual network (ResNet) of 50 layers. In addition to the importance of feature extraction, a machine learning algorithm needs also to be insensitive to outliers while being able to model the context. To tackle this problem, Long Short-Term Memory (LSTM) networks are utilized. The system is then trained in an end-to-end fashion where - by also taking advantage of the correlations of the each of the streams - we manage to significantly outperform the traditional approaches based on auditory and visual handcrafted features for the prediction of spontaneous and natural emotions on the RECOLA database of the AVEC 2016 research challenge on emotion recognition.
References in corpus (3)
Cited by in corpus (57)
- Deep Affect Prediction in-the-wild: Aff-Wild Database and Challenge, Deep Architectures, and Beyond
- Deep Learning for Human Affect Recognition: Insights and New Developments
- Leveraging Recent Advances in Deep Learning for Audio-Visual Emotion Recognition
- Unlocking the Emotional World of Visual Media: An Overview of the Science, Research, and Impact of Understanding Emotion
- Semi-supervised Deep Generative Modelling of Incomplete Multi-Modality Emotional Data
- HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition
- Emotion recognition by fusing time synchronous and time asynchronous representations
- EmoCo: Visual Analysis of Emotion Coherence in Presentation Videos
- Deep Feature Space: A Geometrical Perspective
- Deep Learning based Emotion Recognition System Using Speech Features and Transcriptions
- EmoBed: Strengthening Monomodal Emotion Recognition via Training with Crossmodal Emotion Embeddings
- Self-supervised Audiovisual Representation Learning for Remote Sensing Data
- Dynamic Difficulty Awareness Training for Continuous Emotion Prediction
- Depression Scale Recognition from Audio, Visual and Text Analysis
- Adversarial Training in Affective Computing and Sentiment Analysis: Recent Advances and Perspectives
- Leveraging TCN and Transformer for effective visual-audio fusion in continuous emotion recognition
- Affect-Driven Modelling of Robot Personality for Collaborative Human-Robot Interactions
- Multimodal Utterance-level Affect Analysis using Visual, Audio and Text Features
- Characterizing Types of Convolution in Deep Convolutional Recurrent Neural Networks for Robust Speech Emotion Recognition
- Deep Covariance Descriptors for Facial Expression Recognition
- Emotion Recognition System from Speech and Visual Information based on Convolutional Neural Networks
- Multi-Task Learning with Auxiliary Speaker Identification for Conversational Emotion Recognition
- End-to-End Deep Fault Tolerant Control
- End2You -- The Imperial Toolkit for Multimodal Profiling by End-to-End Learning
- Multi-modal embeddings using multi-task learning for emotion recognition
- Multistage linguistic conditioning of convolutional layers for speech emotion recognition
- Supervised Contrastive Learning for Affect Modelling
- A Transfer Learning Method for Speech Emotion Recognition from Automatic Speech Recognition
- The Next Big Thing(s) in Unsupervised Machine Learning: Five Lessons from Infant Learning
- Convolutional Attention Networks for Multimodal Emotion Recognition from Speech and Text Data
- Estimating the Uncertainty in Emotion Class Labels with Utterance-Specific Dirichlet Priors
- Transfer Learning From Sound Representations For Anger Detection in Speech
- Focal Loss based Residual Convolutional Neural Network for Speech Emotion Recognition
- From the Lab to the Wild: Affect Modeling via Privileged Information
- Dyadic Speech-based Affect Recognition using DAMI-P2C Parent-child Multimodal Interaction Dataset
- Maximum Likelihood Estimation for Multimodal Learning with Missing Modality
- Attention-Augmented End-to-End Multi-Task Learning for Emotion Prediction from Speech
- MuSe 2020 -- The First International Multimodal Sentiment Analysis in Real-life Media Challenge and Workshop
- Single-Channel Speech Separation with Auxiliary Speaker Embeddings
- Snore-GANs: Improving Automatic Snore Sound Classification with Synthesized Data
- Multi-Modal Emotion Detection with Transfer Learning
- Multimodal Continuous Emotion Recognition using Deep Multi-Task Learning with Correlation Loss
- End-to-end Multimodal Emotion and Gender Recognition with Dynamic Joint Loss Weights
- Audio-Visual Kinship Verification
- Towards Robust Deep Neural Networks for Affect and Depression Recognition from Speech
- Multimodal Embeddings from Language Models
- Detecting Escalation Level from Speech with Transfer Learning and Acoustic-Lexical Information Fusion
- Multimodal Knowledge Expansion
- Multimodal End-to-End Group Emotion Recognition using Cross-Modal Attention
- Cross-Modal Knowledge Transfer via Inter-Modal Translation and Alignment for Affect Recognition
- Towards Privacy-Preserving Affect Recognition: A Two-Level Deep Learning Architecture
- Affect-DML: Context-Aware One-Shot Recognition of Human Affect using Deep Metric Learning
- Teaching AI to Feel: A Collaborative, Full-Body Exploration of Emotive Communication
- ROSbag-based Multimodal Affective Dataset for Emotional and Cognitive States
- Acted vs. Improvised: Domain Adaptation for Elicitation Approaches in Audio-Visual Emotion Recognition
- Facial Emotion Recognition using Deep Residual Networks in Real-World Environments
- Toward Multimodal Modeling of Emotional Expressiveness