A vector quantized masked autoencoder for audiovisual speech emotion recognition
arXiv:2305.03568 · doi:10.1016/j.cviu.2025.104362
Abstract
An important challenge in emotion recognition is to develop methods that can leverage unlabeled training data. In this paper, we propose the VQ-MAE-AV model, a self-supervised multimodal model that leverages masked autoencoders to learn representations of audiovisual speech without labels. The model includes vector quantized variational autoencoders that compress raw audio and visual speech data into discrete tokens. The audiovisual speech tokens are used to train a multimodal masked autoencoder that consists of an encoder-decoder architecture with attention mechanisms. The model is designed to extract both local (i.e., at the frame level) and global (i.e., at the sequence level) representations of audiovisual speech. During self-supervised pre-training, the VQ-MAE-AV model is trained on a large-scale unlabeled dataset of audiovisual speech, for the task of reconstructing randomly masked audiovisual speech tokens and with a contrastive learning strategy. During this pre-training, the encoder learns to extract a representation of audiovisual speech that can be subsequently leveraged for emotion recognition. During the supervised fine-tuning stage, a small classification model is trained on top of the VQ-MAE-AV encoder for an emotion recognition task. The proposed approach achieves state-of-the-art emotion recognition results across several datasets in both controlled and in-the-wild conditions.
13 pages, 6 figures, https://samsad35.github.io/VQ-MAE-AudioVisual/
References in corpus (23)
- A Simple Framework for Contrastive Learning of Visual Representations
- Neural Discrete Representation Learning
- Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
- Improved Baselines with Momentum Contrastive Learning
- Unsupervised Representation Learning by Predicting Image Rotations
- How far are we from solving the 2D & 3D Face Alignment problem? (and a dataset of 230,000 3D facial landmarks)
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- Masked Autoencoders As Spatiotemporal Learners
- data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language
- Leveraging Recent Advances in Deep Learning for Audio-Visual Emotion Recognition
- Self-Supervised MultiModal Versatile Networks
- Query2Label: A Simple Transformer Way to Multi-Label Classification
- Masked Autoencoders that Listen
- Aff-Wild2: Extending the Aff-Wild Database for Affect Recognition
- A Fine-tuned Wav2vec 2.0/HuBERT Benchmark For Speech Emotion Recognition, Speaker Verification and Spoken Language Understanding
- A Survey on Masked Autoencoder for Self-supervised Learning in Vision and Beyond
- Augmenting Convolutional networks with attention-based aggregation
- Multimodal Masked Autoencoders Learn Transferable Representations
- A cross-modal fusion network based on self-attention and residual structure for multimodal emotion recognition
- A Multi-modal and Multi-task Learning Method for Action Unit and Expression Recognition
- A multimodal dynamical variational autoencoder for audiovisual speech representation learning
- Emotion Recognition for In-the-wild Videos