Self-Supervised Learning for Videos: A Survey
arXiv:2207.00419 · doi:10.1145/3577925
Abstract
The remarkable success of deep learning in various domains relies on the availability of large-scale annotated datasets. However, obtaining annotations is expensive and requires great effort, which is especially challenging for videos. Moreover, the use of human-generated annotations leads to models with biased learning and poor domain generalization and robustness. As an alternative, self-supervised learning provides a way for representation learning which does not require annotations and has shown promise in both image and video domains. Different from the image domain, learning video representations are more challenging due to the temporal dimension, bringing in motion and other environmental dynamics. This also provides opportunities for video-exclusive ideas that advance self-supervised learning in the video and multimodal domain. In this survey, we provide a review of existing approaches on self-supervised learning focusing on the video domain. We summarize these methods into four different categories based on their learning objectives: 1) pretext tasks, 2) generative learning, 3) contrastive learning, and 4) cross-modal agreement. We further introduce the commonly used datasets, downstream evaluation tasks, insights into the limitations of existing works, and the potential future directions in this area.
ACM CSUR (December 2022). Project Link: https://bit.ly/3Oimc7Q
References in corpus (9)
- Bootstrap your own latent: A new approach to self-supervised Learning
- The Kinetics Human Action Video Dataset
- Generating Videos with Scene Dynamics
- One Model To Learn Them All
- Self-supervised Video Representation Learning Using Inter-intra Contrastive Framework
- The AVA-Kinetics Localized Human Actions Video Dataset
- VideoMoCo: Contrastive Video Representation Learning with Temporally Adversarial Examples
- Representation Learning with Video Deep InfoMax
- Self-supervised Video Representation Learning with Motion-Aware Masked Autoencoders
Cited by in corpus (8)
- Video Transformers: A Survey
- From Google Gemini to OpenAI Q* (Q-Star): A Survey of Reshaping the Generative Artificial Intelligence (AI) Research Landscape
- Self-supervised visual learning in the low-data regime: a comparative evaluation
- Augmentation-aware Self-supervised Learning with Conditioned Projector
- A Hierarchical Framework with Spatio-Temporal Consistency Learning for Emergence Detection in Complex Adaptive Systems
- The ParlaSpeech Collection of Automatically Generated Speech and Text Datasets from Parliamentary Proceedings
- Leveraging Self-Supervised Learning for Fetal Cardiac Planes Classification using Ultrasound Scan Videos
- Advancing Video Self-Supervised Learning via Image Foundation Models