Unsupervised Visual Representation Learning by Context Prediction
arXiv:1505.05192
Abstract
This work explores the use of spatial context as a source of free and plentiful supervisory signal for training a rich visual representation. Given only a large, unlabeled image collection, we extract random pairs of patches from each image and train a convolutional neural net to predict the position of the second patch relative to the first. We argue that doing well on this task requires the model to learn to recognize objects and their parts. We demonstrate that the feature representation learned using this within-image context indeed captures visual similarity across images. For example, this representation allows us to perform unsupervised visual discovery of objects like cats, people, and even birds from the Pascal VOC 2011 detection dataset. Furthermore, we show that the learned ConvNet can be used in the R-CNN framework and provides a significant boost over a randomly-initialized ConvNet, resulting in state-of-the-art performance among algorithms which use only Pascal-provided training set annotations.
Oral paper at ICCV 2015
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Fully Convolutional Networks for Semantic Segmentation
- Unsupervised Learning of Visual Representations using Videos
- Data-dependent Initializations of Convolutional Neural Networks
- Analyzing the Performance of Multilayer Neural Networks for Object Recognition
Cited by in corpus (96)
- The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
- Adversarial Feature Learning
- WebVision Database: Visual Learning and Understanding from Web Data
- What makes ImageNet good for transfer learning?
- Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
- Unsupervised Learning of Visual Representations using Videos
- Data-dependent Initializations of Convolutional Neural Networks
- Unsupervised Feature Learning via Non-Parametric Instance-level Discrimination
- Improvements to context based self-supervised learning
- Self-supervised learning of a facial attribute embedding from video
- Self-Supervised Multisensor Change Detection
- Rethinking ImageNet Pre-training
- Augmenting Supervised Neural Networks with Unsupervised Objectives for Large-scale Image Classification
- PixelNet: Representation of the pixels, by the pixels, and for the pixels
- Loss is its own Reward: Self-Supervision for Reinforcement Learning
- Unsupervised Learning of View-invariant Action Representations
- Learning visual groups from co-occurrences in space and time
- Learning by Association - A versatile semi-supervised training method for neural networks
- Colorization as a Proxy Task for Visual Understanding
- Learning From Noisy Large-Scale Datasets With Minimal Supervision
- Less is More: Surgical Phase Recognition with Less Annotations through Self-Supervised Pre-training of CNN-LSTM Networks
- Time-Contrastive Networks: Self-Supervised Learning from Video
- Patch2Self: Denoising Diffusion MRI with Self-Supervised Learning
- Look, Listen and Learn
- Identifying and Categorizing Anomalies in Retinal Imaging Data
- Unsupervised Learning of Long-Term Motion Dynamics for Videos
- Unsupervised Representation Learning by Sorting Sequences
- Self-Supervised Video Representation Learning with Space-Time Cubic Puzzles
- Learning Image Representations by Completing Damaged Jigsaw Puzzles
- Unsupervised temporal context learning using convolutional neural networks for laparoscopic workflow analysis
- Unsupervised High-level Feature Learning by Ensemble Projection for Semi-supervised Image Classification and Image Clustering
- Leveraging Unlabeled Data for Crowd Counting by Learning to Rank
- Advancing from Predictive Maintenance to Intelligent Maintenance with AI and IIoT
- Self-supervised Scale Equivariant Network for Weakly Supervised Semantic Segmentation
- Transitive Invariance for Self-supervised Visual Representation Learning
- Walk and Learn: Facial Attribute Representation Learning from Egocentric Video and Contextual Data
- An Adversarial Regularisation for Semi-Supervised Training of Structured Output Neural Networks
- Self-supervised learning of visual features through embedding images into text topic spaces
- Scene Parsing with Global Context Embedding
- Emotion Recognition in Speech using Cross-Modal Transfer in the Wild
- Associative Compression Networks for Representation Learning
- WebVision Challenge: Visual Learning and Understanding With Web Data
- Unsupervised Learning from Continuous Video in a Scalable Predictive Recurrent Network
- Split-Brain Autoencoders: Unsupervised Learning by Cross-Channel Prediction
- Skip-Clip: Self-Supervised Spatiotemporal Representation Learning by Future Clip Order Ranking
- Exploring Instance Relations for Unsupervised Feature Embedding
- Exploring Self-Supervised Regularization for Supervised and Semi-Supervised Learning
- Understand and Improve Contrastive Learning Methods for Visual Representation: A Review
- Self-Supervised Feature Learning by Learning to Spot Artifacts
- Unsupervised Single Image Deraining with Self-supervised Constraints
- Unsupervised Joint Mining of Deep Features and Image Labels for Large-scale Radiology Image Categorization and Scene Recognition
- Self-Supervised GANs with Label Augmentation
- On The State of Data In Computer Vision: Human Annotations Remain Indispensable for Developing Deep Learning Models
- Self-Supervised Video Representation Learning With Odd-One-Out Networks
- Revisiting Image Aesthetic Assessment via Self-Supervised Feature Learning
- Unsupervised Learning of Edges
- Unsupervised Vision-and-Language Pre-training Without Parallel Images and Captions
- Auxiliary Tasks Speed Up Learning PointGoal Navigation
- Mix-and-Match Tuning for Self-Supervised Semantic Segmentation
- Spatial Ensemble: a Novel Model Smoothing Mechanism for Student-Teacher Framework
- Probing Emergent Semantics in Predictive Agents via Question Answering
- Video Jigsaw: Unsupervised Learning of Spatiotemporal Context for Video Action Recognition
- Conditional independence for pretext task selection in Self-supervised speech representation learning
- TextTopicNet - Self-Supervised Learning of Visual Features Through Embedding Images on Semantic Text Spaces
- Learning Deep Representations Using Convolutional Auto-encoders with Symmetric Skip Connections
- Cross-Domain Self-supervised Multi-task Feature Learning using Synthetic Imagery
- Motion2Vec: Semi-Supervised Representation Learning from Surgical Videos
- Evolutionary Augmentation Policy Optimization for Self-supervised Learning
- Fine-grained Anomaly Detection via Multi-task Self-Supervision
- On Machine Learning and Structure for Mobile Robots
- Sparse Autoencoder for Unsupervised Nucleus Detection and Representation in Histopathology Images
- Quad-networks: unsupervised learning to rank for interest point detection
- Towards Fine-grained Visual Representations by Combining Contrastive Learning with Image Reconstruction and Attention-weighted Pooling
- Temporal Relational Modeling with Self-Supervision for Action Segmentation
- StackMix: A complementary Mix algorithm
- Ensemble Manifold Segmentation for Model Distillation and Semi-supervised Learning
- Self-Supervised Visual Representations for Cross-Modal Retrieval
- Functional Regularization for Representation Learning: A Unified Theoretical Perspective
- Pseudo-Representation Labeling Semi-Supervised Learning
- Class Subset Selection for Transfer Learning using Submodularity
- DTG-Net: Differentiated Teachers Guided Self-Supervised Video Action Recognition
- The emergence of visual semantics through communication games
- Unsupervised Learning of Face Representations
- Perceive Where to Focus: Learning Visibility-aware Part-level Features for Partial Person Re-identification
- Towards Robust Pattern Recognition: A Review
- Image Reassembly Combining Deep Learning and Shortest Path Problem
- From Same Photo: Cheating on Visual Kinship Challenges
- Laplacian Denoising Autoencoder
- Reliable Label Bootstrapping for Semi-Supervised Learning
- Weakly-Supervised Spatial Context Networks
- Cognitive Inference of Demographic Data by User Ratings
- A Classification approach towards Unsupervised Learning of Visual Representations
- Ridge Rider: Finding Diverse Solutions by Following Eigenvectors of the Hessian
- Actor and Observer: Joint Modeling of First and Third-Person Videos
- Learning to Estimate Pose by Watching Videos
- Unsupervised Holistic Image Generation from Key Local Patches