Learning Features by Watching Objects Move
arXiv:1612.06370
Abstract
This paper presents a novel yet intuitive approach to unsupervised feature learning. Inspired by the human visual system, we explore whether low-level motion-based grouping cues can be used to learn an effective visual representation. Specifically, we use unsupervised motion-based segmentation on videos to obtain segments, which we use as 'pseudo ground truth' to train a convolutional network to segment objects from a single frame. Given the extensive evidence that motion plays a key role in the development of the human visual system, we hope that this straightforward approach to unsupervised learning will be more effective than cleverly designed 'pretext' tasks studied in the literature. Indeed, our extensive experiments show that this is the case. When used for transfer learning on object detection, our representation significantly outperforms previous unsupervised approaches across multiple settings, especially when training data for the target task is scarce.
CVPR 2017
Cited by in corpus (13)
- Unsupervised Learning of Depth and Ego-Motion from Video
- Rethinking ImageNet Pre-training
- Colorization as a Proxy Task for Visual Understanding
- Time-Contrastive Networks: Self-Supervised Learning from Video
- Lucid Data Dreaming for Video Object Segmentation
- Self-Supervised Feature Learning by Learning to Spot Artifacts
- Object Discovery in Videos as Foreground Motion Clustering
- Visual Forecasting by Imitating Dynamics in Natural Sequences
- Mix-and-Match Tuning for Self-Supervised Semantic Segmentation
- Cross-Domain Self-supervised Multi-task Feature Learning using Synthetic Imagery
- See More, Know More: Unsupervised Video Object Segmentation with Co-Attention Siamese Networks
- Unsupervised Feature Learning of Human Actions as Trajectories in Pose Embedding Manifold
- Exploit Clues from Views: Self-Supervised and Regularized Learning for Multiview Object Recognition