Self-Supervised Spatiotemporal Feature Learning via Video Rotation Prediction
arXiv:1811.11387
Abstract
The success of deep neural networks generally requires a vast amount of training data to be labeled, which is expensive and unfeasible in scale, especially for video collections. To alleviate this problem, in this paper, we propose 3DRotNet: a fully self-supervised approach to learn spatiotemporal features from unlabeled videos. A set of rotations are applied to all videos, and a pretext task is defined as prediction of these rotations. When accomplishing this task, 3DRotNet is actually trained to understand the semantic concepts and motions in videos. In other words, it learns a spatiotemporal video representation, which can be transferred to improve video understanding tasks in small datasets. Our extensive experiments successfully demonstrate the effectiveness of the proposed framework on action recognition, leading to significant improvements over the state-of-the-art self-supervised methods. With the self-supervised pre-trained 3DRotNet from large datasets, the recognition accuracy is boosted up by 20.4% on UCF101 and 16.7% on HMDB51 respectively, compared to the models trained from scratch.
References in corpus (7)
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- The Kinetics Human Action Video Dataset
- Decomposing Motion and Content for Natural Video Sequence Prediction
- Temporal 3D ConvNets: New Architecture and Transfer Learning for Video Classification
- Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey
- Self-Supervised Video Representation Learning with Space-Time Cubic Puzzles
- Pose from Action: Unsupervised Learning of Pose Features based on Motion
Cited by in corpus (56)
- Transformers in Vision: A Survey
- Data-Efficient Image Recognition with Contrastive Predictive Coding
- Self-supervised ECG Representation Learning for Emotion Recognition
- Self-supervised Co-training for Video Representation Learning
- Self-Supervised Learning by Cross-Modal Audio-Video Clustering
- Learning Video Representations using Contrastive Bidirectional Transformer
- Self-Supervised MultiModal Versatile Networks
- Self-supervised Visual Feature Learning with Deep Neural Networks: A Survey
- TCLR: Temporal Contrastive Learning for Video Representation
- Audiovisual SlowFast Networks for Video Recognition
- A Comprehensive Study of Deep Video Action Recognition
- Self-supervised remote sensing feature learning: Learning Paradigms, Challenges, and Future Works
- Spatiotemporal Contrastive Video Representation Learning
- Memory-augmented Dense Predictive Coding for Video Representation Learning
- Labelling unlabelled videos from scratch with multi-modal self-supervision
- Video Representation Learning with Visual Tempo Consistency
- Self-supervised Learning for Video Correspondence Flow
- Video Representation Learning by Dense Predictive Coding
- Parameter Efficient Multimodal Transformers for Video Representation Learning
- Can Temporal Information Help with Contrastive Self-Supervised Learning?
- Pose-guided Visible Part Matching for Occluded Person ReID
- Video Understanding as Machine Translation
- VideoMoCo: Contrastive Video Representation Learning with Temporally Adversarial Examples
- Self-supervised Human Activity Recognition by Learning to Predict Cross-Dimensional Motion
- Dual Contrastive Learning for Spatio-temporal Representation
- Learning Video Representations from Textual Web Supervision
- Skip-Clip: Self-Supervised Spatiotemporal Representation Learning by Future Clip Order Ranking
- SnapshotNet: Self-supervised Feature Learning for Point Cloud Data Segmentation Using Minimal Labeled Data
- Self-Supervised Representation Learning for Detection of ACL Tear Injury in Knee MR Videos
- Removing the Background by Adding the Background: Towards Background Robust Self-supervised Video Representation Learning
- Unsupervised Learning from Video with Deep Neural Embeddings
- Video 3D Sampling for Self-supervised Representation Learning
- Pretext-Contrastive Learning: Toward Good Practices in Self-supervised Video Representation Leaning
- MAST: A Memory-Augmented Self-supervised Tracker
- Unsupervised Feature Learning for Point Cloud by Contrasting and Clustering With Graph Convolutional Neural Network
- Self-supervised Motion Learning from Static Images
- Self-supervised Feature Learning by Cross-modality and Cross-view Correspondences
- Self-Supervised Pillar Motion Learning for Autonomous Driving
- ASCNet: Self-supervised Video Representation Learning with Appearance-Speed Consistency
- Self-supervised Video Representation Learning by Context and Motion Decoupling
- Back to the Future: Cycle Encoding Prediction for Self-supervised Contrastive Video Representation Learning
- Towards a Hypothesis on Visual Transformation based Self-Supervision
- Hop-Count Based Self-Supervised Anomaly Detection on Attributed Networks
- Motion-aware Contrastive Video Representation Learning via Foreground-background Merging
- Self-Supervised Video Representation Learning with Meta-Contrastive Network
- Unsupervised Visual Representation Learning by Tracking Patches in Video
- ParamCrop: Parametric Cubic Cropping for Video Contrastive Learning
- Video Contrastive Learning with Global Context
- Unsupervised Few-Shot Action Recognition via Action-Appearance Aligned Meta-Adaptation
- Long Short View Feature Decomposition via Contrastive Video Representation Learning
- ArrowGAN : Learning to Generate Videos by Learning Arrow of Time
- Contrastive Learning of Image Representations with Cross-Video Cycle-Consistency
- MIO : Mutual Information Optimization using Self-Supervised Binary Contrastive Learning
- ScaleNet: An Unsupervised Representation Learning Method for Limited Information
- Adaptive and Iteratively Improving Recurrent Lateral Connections
- Self-supervision of Feature Transformation for Further Improving Supervised Learning