Video Transformers: A Survey
arXiv:2201.05991 · doi:10.1109/TPAMI.2023.3243465
Abstract
Transformer models have shown great success handling long-range interactions, making them a promising tool for modeling video. However, they lack inductive biases and scale quadratically with input length. These limitations are further exacerbated when dealing with the high dimensionality introduced by the temporal dimension. While there are surveys analyzing the advances of Transformers for vision, none focus on an in-depth analysis of video-specific designs. In this survey, we analyze the main contributions and trends of works leveraging Transformers to model video. Specifically, we delve into how videos are handled at the input level first. Then, we study the architectural changes made to deal with video more efficiently, reduce redundancy, re-introduce useful inductive biases, and capture long-term temporal dynamics. In addition, we provide an overview of different training regimes and explore effective self-supervised learning strategies for video. Finally, we conduct a performance comparison on the most common benchmark for Video Transformers (i.e., action classification), finding them to outperform 3D ConvNets even with less computational complexity.
References in corpus (23)
- Bootstrap your own latent: A new approach to self-supervised Learning
- Is Space-Time Attention All You Need for Video Understanding?
- Zero-Shot Text-to-Image Generation
- Barlow Twins: Self-Supervised Learning via Redundancy Reduction
- CoAtNet: Marrying Convolution and Attention for All Data Sizes
- Attention is not Explanation
- VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
- VATT: Transformers for Multimodal Self-Supervised Learning from Raw Video, Audio and Text
- Intriguing Properties of Vision Transformers
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning
- Self-Supervised Learning for Videos: A Survey
- Thinking Fast and Slow: Efficient Text-to-Visual Retrieval with Transformers
- Demystifying Contrastive Self-Supervised Learning: Invariances, Augmentations and Dataset Biases
- When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations
- Referring Segmentation in Images and Videos with Cross-Modal Self-Attention Network
- COOT: Cooperative Hierarchical Transformer for Video-Text Representation Learning
- Video Instance Segmentation using Inter-Frame Communication Transformers
- Space-time Mixing Attention for Video Transformer
- A Survey of Visual Transformers
- Deepfake Detection Scheme Based on Vision Transformer and Distillation
- Co-training Transformer with Videos and Images Improves Action Recognition
- Transformers Meet Visual Learning Understanding: A Comprehensive Review
- Time Is MattEr: Temporal Self-supervision for Video Transformers
Cited by in corpus (9)
- Reviewing Intelligent Cinematography: AI research for camera-based video production
- Eat-Radar: Continuous Fine-Grained Intake Gesture Detection Using FMCW Radar and 3D Temporal Convolutional Network with Attention
- Advances in Artificial Intelligence: A Review for the Creative Industries
- Artificial Inductive Bias for Synthetic Tabular Data Generation in Data-Scarce Scenarios
- DiVa-360: The Dynamic Visual Dataset for Immersive Neural Fields
- Reversing the Damage: A QP-Aware Transformer-Diffusion Approach for 8K Video Restoration under Codec Compression
- Video Flow as Time Series: Discovering Temporal Consistency and Variability for VideoQA
- The Unreasonable Effectiveness of Large Language-Vision Models for Source-free Video Domain Adaptation
- An Overview on Generative AI at Scale with Edge-Cloud Computing