Self-supervised Motion Learning from Static Images
arXiv:2104.00240
Abstract
Motions are reflected in videos as the movement of pixels, and actions are essentially patterns of inconsistent motions between the foreground and the background. To well distinguish the actions, especially those with complicated spatio-temporal interactions, correctly locating the prominent motion areas is of crucial importance. However, most motion information in existing videos are difficult to label and training a model with good motion representations with supervision will thus require a large amount of human labour for annotation. In this paper, we address this problem by self-supervised learning. Specifically, we propose to learn Motion from Static Images (MoSI). The model learns to encode motion information by classifying pseudo motions generated by MoSI. We furthermore introduce a static mask in pseudo motions to create local motion patterns, which forces the model to additionally locate notable motion areas for the correct classification.We demonstrate that MoSI can discover regions with large motion even without fine-tuning on the downstream datasets. As a result, the learned motion representations boost the performance of tasks requiring understanding of complex scenes and motions, i.e., action recognition. Extensive experiments show the consistent and transferable improvements achieved by MoSI. Codes will be soon released.
To appear in CVPR 2021
References in corpus (5)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
- The Kinetics Human Action Video Dataset
- Memory-augmented Dense Predictive Coding for Video Representation Learning
- Video Cloze Procedure for Self-Supervised Spatio-Temporal Learning
Cited by in corpus (6)
- Proposal Relation Network for Temporal Action Detection
- Relation Modeling in Spatio-Temporal Action Localization
- Towards Training Stronger Video Vision Transformers for EPIC-KITCHENS-100 Action Recognition
- Exploring Stronger Feature for Temporal Action Localization
- Weakly-Supervised Temporal Action Localization Through Local-Global Background Modeling
- A Stronger Baseline for Ego-Centric Action Detection