A Closer Look at Spatiotemporal Convolutions for Action Recognition
arXiv:1711.11248
Abstract
In this paper we discuss several forms of spatiotemporal convolutions for video analysis and study their effects on action recognition. Our motivation stems from the observation that 2D CNNs applied to individual frames of the video have remained solid performers in action recognition. In this work we empirically demonstrate the accuracy advantages of 3D CNNs over 2D CNNs within the framework of residual learning. Furthermore, we show that factorizing the 3D convolutional filters into separate spatial and temporal components yields significantly advantages in accuracy. Our empirical study leads to the design of a new spatiotemporal convolutional block "R(2+1)D" which gives rise to CNNs that achieve results comparable or superior to the state-of-the-art on Sports-1M, Kinetics, UCF101 and HMDB51.
References in corpus (4)
Cited by in corpus (12)
- Billion-scale semi-supervised learning for image classification
- Self-Supervised Learning for Videos: A Survey
- TA2N: Two-Stage Action Alignment Network for Few-shot Action Recognition
- Graph-Based Global Reasoning Networks
- Multi-Fiber Networks for Video Recognition
- Enhanced 3D convolutional networks for crowd counting
- Spatio-temporal Action Recognition: A Survey
- MotionSqueeze: Neural Motion Feature Learning for Video Understanding
- Tracking Without Re-recognition in Humans and Machines
- Action-Based ADHD Diagnosis in Video
- Image to Video Domain Adaptation Using Web Supervision
- Efficient Modelling Across Time of Human Actions and Interactions