Unsupervised Learning of Visual Representations by Solving Jigsaw Puzzles
arXiv:1603.09246
Abstract
In this paper we study the problem of image representation learning without human annotation. By following the principles of self-supervision, we build a convolutional neural network (CNN) that can be trained to solve Jigsaw puzzles as a pretext task, which requires no manual labeling, and then later repurposed to solve object classification and detection. To maintain the compatibility across tasks we introduce the context-free network (CFN), a siamese-ennead CNN. The CFN takes image tiles as input and explicitly limits the receptive field (or context) of its early processing units to one tile at a time. We show that the CFN includes fewer parameters than AlexNet while preserving the same semantic learning capabilities. By training the CFN to solve Jigsaw puzzles, we learn both a feature mapping of object parts as well as their correct spatial arrangement. Our experimental evaluations show that the learned features capture semantically relevant content. Our proposed method for learning visual representations outperforms state of the art methods in several transfer learning benchmarks.
ECCV 2016
References in corpus (5)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Two-Stream Convolutional Networks for Action Recognition in Videos
- How transferable are features in deep neural networks?
- Understanding Neural Networks Through Deep Visualization
- Data-dependent Initializations of Convolutional Neural Networks
Cited by in corpus (6)
- Deep Successor Reinforcement Learning
- Colorful Image Colorization
- Cross-domain few-shot learning with unlabelled data
- Quad-networks: unsupervised learning to rank for interest point detection
- Exploiting Spatio-Temporal Structure with Recurrent Winner-Take-All Networks
- About Explicit Variance Minimization: Training Neural Networks for Medical Imaging With Limited Data Annotations