SVFAP: Self-supervised Video Facial Affect Perceiver
arXiv:2401.00416 · doi:10.1109/TAFFC.2024.3436913
Abstract
Video-based facial affect analysis has recently attracted increasing attention owing to its critical role in human-computer interaction. Previous studies mainly focus on developing various deep learning architectures and training them in a fully supervised manner. Although significant progress has been achieved by these supervised methods, the longstanding lack of large-scale high-quality labeled data severely hinders their further improvements. Motivated by the recent success of self-supervised learning in computer vision, this paper introduces a self-supervised approach, termed Self-supervised Video Facial Affect Perceiver (SVFAP), to address the dilemma faced by supervised methods. Specifically, SVFAP leverages masked facial video autoencoding to perform self-supervised pre-training on massive unlabeled facial videos. Considering that large spatiotemporal redundancy exists in facial videos, we propose a novel temporal pyramid and spatial bottleneck Transformer as the encoder of SVFAP, which not only largely reduces computational costs but also achieves excellent performance. To verify the effectiveness of our method, we conduct experiments on nine datasets spanning three downstream tasks, including dynamic facial expression recognition, dimensional emotion recognition, and personality recognition. Comprehensive results demonstrate that SVFAP can learn powerful affect-related representations via large-scale self-supervised pre-training and it significantly outperforms previous state-of-the-art methods on all datasets. Code is available at https://github.com/sunlicai/SVFAP.
Published in: IEEE Transactions on Affective Computing (Early Access). The code and models are available at https://github.com/sunlicai/SVFAP
References in corpus (20)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Gaussian Error Linear Units (GELUs)
- VoxCeleb2: Deep Speaker Recognition
- Deep Facial Expression Recognition: A Survey
- Is Space-Time Attention All You Need for Video Understanding?
- Self-supervised Learning: Generative or Contrastive
- BEiT: BERT Pre-Training of Image Transformers
- Deep Learning for Human Affect Recognition: Insights and New Developments
- Masked Autoencoders As Spatiotemporal Learners
- Self-supervised learning of a facial attribute embedding from video
- Deep Impression: Audiovisual Deep Residual Networks for Multimodal Apparent Personality Trait Recognition
- PersEmoN: A Deep Network for Joint Analysis of Apparent Personality, Emotion and Their Relationship
- MSAF: Multimodal Split Attention Fusion
- A Survey on Masked Autoencoder for Self-supervised Learning in Vision and Beyond
- The Ambiguous World of Emotion Representation
- Spatio-Temporal Transformer for Dynamic Facial Expression Recognition in the Wild
- A cross-modal fusion network based on self-attention and residual structure for multimodal emotion recognition
- NR-DFERNet: Noise-Robust Network for Dynamic Facial Expression Recognition
Cited by in corpus (3)
- HiCMAE: Hierarchical Contrastive Masked Autoencoder for Self-Supervised Audio-Visual Emotion Recognition
- Static for Dynamic: Towards a Deeper Understanding of Dynamic Facial Expressions Using Static Expression Data
- Soften the Mask: Adaptive Temporal Soft Mask for Efficient Dynamic Facial Expression Recognition