Multimodal End-to-End Sparse Model for Emotion Recognition
arXiv:2103.09666
Abstract
Existing works on multimodal affective computing tasks, such as emotion recognition, generally adopt a two-phase pipeline, first extracting feature representations for each single modality with hand-crafted algorithms and then performing end-to-end learning with the extracted features. However, the extracted features are fixed and cannot be further fine-tuned on different target tasks, and manually finding feature extraction algorithms does not generalize or scale well to different tasks, which can lead to sub-optimal performance. In this paper, we develop a fully end-to-end model that connects the two phases and optimizes them jointly. In addition, we restructure the current datasets to enable the fully end-to-end training. Furthermore, to reduce the computational overhead brought by the end-to-end model, we introduce a sparse cross-modal attention mechanism for the feature extraction. Experimental results show that our fully end-to-end model significantly surpasses the current state-of-the-art models based on the two-phase pipeline. Moreover, by adding the sparse cross-modal attention, our model can maintain performance with around half the computation in the feature extraction part.
12 pages, 6 figures
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- VisualBERT: A Simple and Performant Baseline for Vision and Language
- The Curious Case of Neural Text Degeneration
- Submanifold Sparse Convolutional Networks
- Learning Factorized Multimodal Representations
- EmoGraph: Capturing Emotion Correlations using Graph Networks
- Multi-task Learning for Multi-modal Emotion Recognition and Sentiment Analysis
- Kungfupanda at SemEval-2020 Task 12: BERT-Based Multi-Task Learning for Offensive Language Detection