T: Multi-Modal Continuous Valence-Arousal Estimation in the Wild
arXiv:2002.02957
Abstract
This report describes a multi-modal multi-task (T) approach underlying our submission to the valence-arousal estimation track of the Affective Behavior Analysis in-the-wild (ABAW) Challenge, held in conjunction with the IEEE International Conference on Automatic Face and Gesture Recognition (FG) 2020. In the proposed T framework, we fuse both visual features from videos and acoustic features from the audio tracks to estimate the valence and arousal. The spatio-temporal visual features are extracted with a 3D convolutional network and a bidirectional recurrent neural network. Considering the correlations between valence / arousal, emotions, and facial actions, we also explores mechanisms to benefit from other tasks. We evaluated the T framework on the validation set provided by ABAW and it significantly outperforms the baseline method.
6 pages, technical report; submission to ABAW Challenge at FG 2020
References in corpus (6)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- RetinaFace: Single-stage Dense Face Localisation in the Wild
- Expression, Affect, Action Unit Recognition: Aff-Wild2, Multi-Task Learning and ArcFace
- Recurrent Neural Network Transducer for Audio-Visual Speech Recognition
- Facial Affect Recognition in the Wild Using Multi-Task Learning Convolutional Network
- Adversarial-based neural networks for affect estimations in the wild