An audiovisual and contextual approach for categorical and continuous emotion recognition in-the-wild
arXiv:2107.03465 · doi:10.1109/ICCVW54120.2021.00407
Abstract
In this work we tackle the task of video-based audio-visual emotion recognition, within the premises of the 2nd Workshop and Competition on Affective Behavior Analysis in-the-wild (ABAW2). Poor illumination conditions, head/body orientation and low image resolution constitute factors that can potentially hinder performance in case of methodologies that solely rely on the extraction and analysis of facial features. In order to alleviate this problem, we leverage both bodily and contextual features, as part of a broader emotion recognition framework. We choose to use a standard CNN-RNN cascade as the backbone of our proposed model for sequence-to-sequence (seq2seq) learning. Apart from learning through the RGB input modality, we construct an aural stream which operates on sequences of extracted mel-spectrograms. Our extensive experiments on the challenging and newly assembled Aff-Wild2 dataset verify the validity of our intuitive multi-stream and multi-modal approach towards emotion recognition in-the-wild. Emphasis is being laid on the the beneficial influence of the human body and scene context, as aspects of the emotion recognition process that have been left relatively unexplored up to this point. All the code was implemented using PyTorch and is publicly available.
7 pages, 1 figure, 3 tables, accepted to the 2nd Workshop and Competition on Affective Behavior Analysis In-the-Wild (ABAW2)
References in corpus (11)
- Context Based Emotion Recognition using EMOTIC Dataset
- Distribution Matching for Heterogeneous Multi-Task Learning: a Large-scale Face Study
- Affect Analysis in-the-wild: Valence-Arousal, Expressions, Action Units and a Unified Framework
- Expression, Affect, Action Unit Recognition: Aff-Wild2, Multi-Task Learning and ArcFace
- Two-Stream Aural-Visual Affect Analysis in the Wild
- Exploiting Emotional Dependencies with Graph Convolutional Networks for Facial Expression Recognition
- Prior Aided Streaming Network for Multi-task Affective Recognitionat the 2nd ABAW2 Competition
- A Multi-modal and Multi-task Learning Method for Action Unit and Expression Recognition
- Leveraging Semantic Scene Characteristics and Multi-Stream Convolutional Architectures in a Contextual Approach for Video-Based Visual Emotion Recognition in the Wild
- Emotion Recognition with Incomplete Labels Using Modified Multi-task Learning Technique
- Emotion Recognition for In-the-wild Videos