Learning from Temporal Gradient for Semi-supervised Action Recognition
arXiv:2111.13241
Abstract
Semi-supervised video action recognition tends to enable deep neural networks to achieve remarkable performance even with very limited labeled data. However, existing methods are mainly transferred from current image-based methods (e.g., FixMatch). Without specifically utilizing the temporal dynamics and inherent multimodal attributes, their results could be suboptimal. To better leverage the encoded temporal information in videos, we introduce temporal gradient as an additional modality for more attentive feature extraction in this paper. To be specific, our method explicitly distills the fine-grained motion representations from temporal gradient (TG) and imposes consistency across different modalities (i.e., RGB and TG). The performance of semi-supervised action recognition is significantly improved without additional computation or parameters during inference. Our method achieves the state-of-the-art performance on three video action recognition benchmarks (i.e., Kinetics-400, UCF-101, and HMDB-51) under several typical semi-supervised settings (i.e., different ratios of labeled data).
CVPR 2022
References in corpus (14)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Distilling the Knowledge in a Neural Network
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- The Kinetics Human Action Video Dataset
- Is Space-Time Attention All You Need for Video Understanding?
- Towards Good Practices for Very Deep Two-Stream ConvNets
- A Short Note on the Kinetics-700-2020 Human Action Dataset
- ConvNet Architecture Search for Spatiotemporal Feature Learning
- Self-supervised Co-training for Video Representation Learning
- PseudoSeg: Designing Pseudo Labels for Semantic Segmentation
- Would Mega-scale Datasets Further Enhance Spatiotemporal 3D CNNs?
- In Defense of Pseudo-Labeling: An Uncertainty-Aware Pseudo-label Selection Framework for Semi-Supervised Learning
- Rethinking "Batch" in BatchNorm