Unsupervised Learning of Visual Representations using Videos
arXiv:1505.00687
Abstract
Is strong supervision necessary for learning a good visual representation? Do we really need millions of semantically-labeled images to train a Convolutional Neural Network (CNN)? In this paper, we present a simple yet surprisingly powerful approach for unsupervised learning of CNN. Specifically, we use hundreds of thousands of unlabeled videos from the web to learn visual representations. Our key idea is that visual tracking provides the supervision. That is, two patches connected by a track should have similar visual representation in deep feature space since they probably belong to the same object or object part. We design a Siamese-triplet network with a ranking loss function to train this CNN representation. Without using a single image from ImageNet, just using 100K unlabeled videos and the VOC 2012 dataset, we train an ensemble of unsupervised networks that achieves 52% mAP (no bounding box regression). This performance comes tantalizingly close to its ImageNet-supervised counterpart, an ensemble which achieves a mAP of 54.4%. We also show that our unsupervised network can perform competitively in other tasks such as surface-normal estimation.
References in corpus (6)
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Unsupervised Visual Representation Learning by Context Prediction
- Designing Deep Networks for Surface Normal Estimation
- Pedestrian Detection with Unsupervised Multi-Stage Feature Learning
- Computational Baby Learning
- Matching-CNN Meets KNN: Quasi-Parametric Human Parsing
Cited by in corpus (122)
- Generating Videos with Scene Dynamics
- The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
- Adversarial Feature Learning
- Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning
- SfM-Net: Learning of Structure and Motion from Video
- Recent Advances in Convolutional Neural Networks
- What makes ImageNet good for transfer learning?
- Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
- Unsupervised Visual Representation Learning by Context Prediction
- Unsupervised Learning of Depth and Ego-Motion from Video
- Data-dependent Initializations of Convolutional Neural Networks
- Unsupervised Feature Learning via Non-Parametric Instance-level Discrimination
- Joint Unsupervised Learning of Deep Representations and Image Clusters
- Improvements to context based self-supervised learning
- Local Similarity-Aware Deep Feature Embedding
- Rethinking Feature Discrimination and Polymerization for Large-scale Recognition
- Self-supervised learning of a facial attribute embedding from video
- Rethinking ImageNet Pre-training
- Augmenting Supervised Neural Networks with Unsupervised Objectives for Large-scale Image Classification
- PixelNet: Representation of the pixels, by the pixels, and for the pixels
- Generalisation and Sharing in Triplet Convnets for Sketch based Visual Search
- Memory-augmented Dense Predictive Coding for Video Representation Learning
- Generic 3D Representation via Pose Estimation and Matching
- Learning Deep Features via Congenerous Cosine Loss for Person Recognition
- Learning From Noisy Large-Scale Datasets With Minimal Supervision
- Colorization as a Proxy Task for Visual Understanding
- Time-Contrastive Networks: Self-Supervised Learning from Video
- CliqueCNN: Deep Unsupervised Exemplar Learning
- Ensembles of Generative Adversarial Networks
- Look, Listen and Learn
- Video2GIF: Automatic Generation of Animated GIFs from Video
- Temporal Generative Adversarial Nets with Singular Value Clipping
- Unsupervised Learning of Long-Term Motion Dynamics for Videos
- SE3-Pose-Nets: Structured Deep Dynamics Models for Visuomotor Planning and Control
- Video Summarization using Deep Semantic Features
- RankIQA: Learning from Rankings for No-reference Image Quality Assessment
- Colorful Image Colorization
- Unsupervised Representation Learning by Sorting Sequences
- Self-Supervised Video Representation Learning with Space-Time Cubic Puzzles
- Learning Image Representations by Completing Damaged Jigsaw Puzzles
- Unsupervised High-level Feature Learning by Ensemble Projection for Semi-supervised Image Classification and Image Clustering
- DiscrimNet: Semi-Supervised Action Recognition from Videos using Generative Adversarial Networks
- Transitive Invariance for Self-supervised Visual Representation Learning
- Learning Latent Plans from Play
- Localizing and Orienting Street Views Using Overhead Imagery
- DarkRank: Accelerating Deep Metric Learning via Cross Sample Similarities Transfer
- Walk and Learn: Facial Attribute Representation Learning from Egocentric Video and Contextual Data
- Unsupervised learning from videos using temporal coherency deep networks
- Few-Shot Image Recognition by Predicting Parameters from Activations
- Self-supervised learning of visual features through embedding images into text topic spaces
- Blazingly Fast Video Object Segmentation with Pixel-Wise Metric Learning
- Dual Motion GAN for Future-Flow Embedded Video Prediction
- Pose from Action: Unsupervised Learning of Pose Features based on Motion
- Self-Supervised Learning for Spinal MRIs
- Scene Parsing with Global Context Embedding
- WebVision Challenge: Visual Learning and Understanding With Web Data
- Memory Warps for Learning Long-Term Online Video Representations
- Split-Brain Autoencoders: Unsupervised Learning by Cross-Channel Prediction
- Temporal Dynamic Graph LSTM for Action-driven Video Object Detection
- Skip-Clip: Self-Supervised Spatiotemporal Representation Learning by Future Clip Order Ranking
- CortexNet: a Generic Network Family for Robust Visual Temporal Representations
- Unsupervised Category Discovery via Looped Deep Pseudo-Task Optimization Using a Large Scale Radiology Image Database
- Class Rectification Hard Mining for Imbalanced Deep Learning
- Hard-Aware Deeply Cascaded Embedding
- Understand and Improve Contrastive Learning Methods for Visual Representation: A Review
- Efficient end-to-end learning for quantizable representations
- Self-Supervised Feature Learning by Learning to Spot Artifacts
- ActionFlowNet: Learning Motion Representation for Action Recognition
- Unsupervised Single Image Deraining with Self-supervised Constraints
- Self-Supervised Relative Depth Learning for Urban Scene Understanding
- Contextual Visual Similarity
- Object-Centric Representation Learning from Unlabeled Videos
- Self-Supervised Video Representation Learning With Odd-One-Out Networks
- Unsupervised Learning of Edges
- Unsupervised Learning for Large-Scale Fiber Detection and Tracking in Microscopic Material Images
- A-Fast-RCNN: Hard Positive Generation via Adversary for Object Detection
- SampleAhead: Online Classifier-Sampler Communication for Learning from Synthesized Data
- Visual Forecasting by Imitating Dynamics in Natural Sequences
- Mix-and-Match Tuning for Self-Supervised Semantic Segmentation
- Unsupervised Learning of Semantic Audio Representations
- Probing Emergent Semantics in Predictive Agents via Question Answering
- Motion2Vec: Semi-Supervised Representation Learning from Surgical Videos
- An Uncertain Future: Forecasting from Static Images using Variational Autoencoders
- TextTopicNet - Self-Supervised Learning of Visual Features Through Embedding Images on Semantic Text Spaces
- Energy Confused Adversarial Metric Learning for Zero-Shot Image Retrieval and Clustering
- Learning Deep Representations Using Convolutional Auto-encoders with Symmetric Skip Connections
- Quad-networks: unsupervised learning to rank for interest point detection
- Cross-Domain Self-supervised Multi-task Feature Learning using Synthetic Imagery
- Imbalanced Deep Learning by Minority Class Incremental Rectification
- Towards Human-Machine Cooperation: Self-supervised Sample Mining for Object Detection
- Disentangling Motion, Foreground and Background Features in Videos
- Attention Transfer from Web Images for Video Recognition
- Generative Image Modeling using Style and Structure Adversarial Networks
- Better and Faster: Knowledge Transfer from Multiple Self-supervised Learning Tasks via Graph Distillation for Video Classification
- From Images to 3D Shape Attributes
- Exploiting Spatio-Temporal Structure with Recurrent Winner-Take-All Networks
- Class Subset Selection for Transfer Learning using Submodularity
- Ensemble Manifold Segmentation for Model Distillation and Semi-supervised Learning
- Time-Aware and View-Aware Video Rendering for Unsupervised Representation Learning
- Self-Supervised Visual Representations for Cross-Modal Retrieval
- Accurate Deep Representation Quantization with Gradient Snapping Layer for Similarity Search
- Unsupervised Learning of Face Representations
- DTG-Net: Differentiated Teachers Guided Self-Supervised Video Action Recognition
- AETv2: AutoEncoding Transformations for Self-Supervised Representation Learning by Minimizing Geodesic Distances in Lie Groups
- Laplacian Denoising Autoencoder
- Towards Robust Pattern Recognition: A Review
- Exploit Clues from Views: Self-Supervised and Regularized Learning for Multiview Object Recognition
- Unsupervisedly Learned Representations: Should the Quest be Over?
- Video Region Annotation with Sparse Bounding Boxes
- Weakly-Supervised Spatial Context Networks
- Large Scale Novel Object Discovery in 3D
- Learning Rich Representations For Structured Visual Prediction Tasks
- Human-In-The-Loop Person Re-Identification
- Multi-Hypothesis Visual-Inertial Flow
- Online Descriptor Enhancement via Self-Labelling Triplets for Visual Data Association
- Improving Deep Binary Embedding Networks by Order-aware Reweighting of Triplets
- A Self-supervised Learning System for Object Detection in Videos Using Random Walks on Graphs
- A Classification approach towards Unsupervised Learning of Visual Representations
- Ambient Sound Provides Supervision for Visual Learning
- Learning Joint Representations of Videos and Sentences with Web Image Search
- Discriminate-and-Rectify Encoders: Learning from Image Transformation Sets
- Representation Learning by Reconstructing Neighborhoods