A Comprehensive Study of Deep Video Action Recognition
arXiv:2012.06567
Abstract
Video action recognition is one of the representative tasks for video understanding. Over the last decade, we have witnessed great advancements in video action recognition thanks to the emergence of deep learning. But we also encountered new challenges, including modeling long-range temporal information in videos, high computation costs, and incomparable results due to datasets and evaluation protocol variances. In this paper, we provide a comprehensive survey of over 200 existing papers on deep learning for video action recognition. We first introduce the 17 video action recognition datasets that influenced the design of models. Then we present video action recognition models in chronological order: starting with early attempts at adapting deep learning, then to the two-stream networks, followed by the adoption of 3D convolutional kernels, and finally to the recent compute-efficient models. In addition, we benchmark popular methods on several representative datasets and release code for reproducibility. In the end, we discuss open problems and shed light on opportunities for video action recognition to facilitate new research ideas.
Technical report. Code and model zoo can be found at https://cv.gluon.ai/model_zoo/action_recognition.html
References in corpus (22)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- The Kinetics Human Action Video Dataset
- Improved Regularization of Convolutional Neural Networks with Cutout
- YouTube-8M: A Large-Scale Video Classification Benchmark
- Spatiotemporal Residual Networks for Video Action Recognition
- Towards Good Practices for Very Deep Two-Stream ConvNets
- Temporal 3D ConvNets: New Architecture and Transfer Learning for Video Classification
- Bag of Visual Words and Fusion Methods for Action Recognition: Comprehensive Study and Good Practice
- Multi-Task Zero-Shot Action Recognition with Prioritised Data Augmentation
- Why Can't I Dance in the Mall? Learning to Mitigate Scene Bias in Action Recognition
- The AVA-Kinetics Localized Human Actions Video Dataset
- Memory-augmented Dense Predictive Coding for Video Representation Learning
- Joint-task Self-supervised Learning for Temporal Correspondence
- PAN: Towards Fast Action Recognition via Learning Persistence of Appearance
- What have we learned from deep representations for action recognition?
- Efficient Two-Stream Motion and Appearance 3D CNNs for Video Classification
- Heuristic Black-box Adversarial Attacks on Video Recognition Models
- Temporal Interlacing Network
- A3D: Adaptive 3D Networks for Video Action Recognition
- Training Kinetics in 15 Minutes: Large-scale Distributed Training on Videos
- The Best of Both Worlds: Combining Data-independent and Data-driven Approaches for Action Recognition
Cited by in corpus (8)
- All One Needs to Know about Metaverse: A Complete Survey on Technological Singularity, Virtual Ecosystem, and Research Agenda
- IntFormer: Predicting pedestrian intention with the aid of the Transformer architecture
- TEAM-Net: Multi-modal Learning for Video Action Recognition with Partial Decoding
- Evaluating Transformers for Lightweight Action Recognition
- Shaping embodied agent behavior with activity-context priors from egocentric video
- Recursive Fusion and Deformable Spatiotemporal Attention for Video Compression Artifact Reduction
- Survey: Transformer based Video-Language Pre-training
- Video Contrastive Learning with Global Context