Learning to See by Moving
arXiv:1505.01596
Abstract
The dominant paradigm for feature learning in computer vision relies on training neural networks for the task of object recognition using millions of hand labelled images. Is it possible to learn useful features for a diverse set of visual tasks using any other form of supervision? In biology, living organisms developed the ability of visual perception for the purpose of moving and acting in the world. Drawing inspiration from this observation, in this work we investigate if the awareness of egomotion can be used as a supervisory signal for feature learning. As opposed to the knowledge of class labels, information about egomotion is freely available to mobile agents. We show that given the same number of training images, features learnt using egomotion as supervision compare favourably to features learnt using class-label as supervision on visual tasks of scene recognition, object recognition, visual odometry and keypoint matching.
12 pages
References in corpus (1)
Cited by in corpus (85)
- Unsupervised Representation Learning by Predicting Image Rotations
- The Unreasonable Effectiveness of Deep Features as a Perceptual Metric
- Adversarial Feature Learning
- Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics
- Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning
- Recent Advances in Convolutional Neural Networks
- Unsupervised CNN for Single View Depth Estimation: Geometry to the Rescue
- WebVision Database: Visual Learning and Understanding from Web Data
- What makes ImageNet good for transfer learning?
- Universal Correspondence Network
- Spatio-temporal video autoencoder with differentiable memory
- Unsupervised Learning of Depth and Ego-Motion from Video
- Data-dependent Initializations of Convolutional Neural Networks
- Improvements to context based self-supervised learning
- Discovery of Latent 3D Keypoints via End-to-end Geometric Reasoning
- Equivariance Through Parameter-Sharing
- Self-supervised learning of a facial attribute embedding from video
- PixelNet: Representation of the pixels, by the pixels, and for the pixels
- Unsupervised Learning of Monocular Depth Estimation and Visual Odometry with Deep Feature Reconstruction
- Memory-augmented Dense Predictive Coding for Video Representation Learning
- Learning visual groups from co-occurrences in space and time
- Generic 3D Representation via Pose Estimation and Matching
- Learning Physical Intuition of Block Towers by Example
- A Framework For Contrastive Self-Supervised Learning And Designing A New Approach
- 3DMatch: Learning Local Geometric Descriptors from RGB-D Reconstructions
- Learning to Perform Physics Experiments via Deep Reinforcement Learning
- DeepVO: A Deep Learning approach for Monocular Visual Odometry
- Look, Listen and Learn
- WarpNet: Weakly Supervised Matching for Single-view Reconstruction
- Learning Dense Correspondence via 3D-guided Cycle Consistency
- Colorful Image Colorization
- Unsupervised Representation Learning by Sorting Sequences
- Learning Image Representations by Completing Damaged Jigsaw Puzzles
- Unsupervised High-level Feature Learning by Ensemble Projection for Semi-supervised Image Classification and Image Clustering
- Transitive Invariance for Self-supervised Visual Representation Learning
- Localizing and Orienting Street Views Using Overhead Imagery
- Walk and Learn: Facial Attribute Representation Learning from Egocentric Video and Contextual Data
- Self-supervised learning of visual features through embedding images into text topic spaces
- Patterns for Learning with Side Information
- Pose from Action: Unsupervised Learning of Pose Features based on Motion
- WebVision Challenge: Visual Learning and Understanding With Web Data
- Split-Brain Autoencoders: Unsupervised Learning by Cross-Channel Prediction
- Unsupervised learning of object frames by dense equivariant image labelling
- ENG: End-to-end Neural Geometry for Robust Depth and Pose Estimation using CNNs
- Automated Top View Registration of Broadcast Football Videos
- Self-Supervised Feature Learning by Learning to Spot Artifacts
- Object-Centric Representation Learning from Unlabeled Videos
- Unsupervised Single Image Deraining with Self-supervised Constraints
- Social Behavior Prediction from First Person Videos
- Self-supervisory Signals for Object Discovery and Detection
- Self-Supervised Relative Depth Learning for Urban Scene Understanding
- Siamese Convolutional Neural Network for Sub-millimeter-accurate Camera Pose Estimation and Visual Servoing
- Object category learning and retrieval with weak supervision
- Self-Supervised Video Representation Learning With Odd-One-Out Networks
- Deep Semantic Architecture with discriminative feature visualization for neuroimage analysis
- Video Jigsaw: Unsupervised Learning of Spatiotemporal Context for Video Action Recognition
- WGANVO: Monocular Visual Odometry based on Generative Adversarial Networks
- Unsupervised Learning of Semantic Audio Representations
- Convolutional Patch Representations for Image Retrieval: an Unsupervised Approach
- TextTopicNet - Self-Supervised Learning of Visual Features Through Embedding Images on Semantic Text Spaces
- Cross-Domain Self-supervised Multi-task Feature Learning using Synthetic Imagery
- An Uncertain Future: Forecasting from Static Images using Variational Autoencoders
- Towards CNN Map Compression for camera relocalisation
- Transferring Physical Motion Between Domains for Neural Inertial Tracking
- Generative Image Modeling using Style and Structure Adversarial Networks
- Exploiting Spatio-Temporal Structure with Recurrent Winner-Take-All Networks
- Class Subset Selection for Transfer Learning using Submodularity
- Self-Supervised Visual Representations for Cross-Modal Retrieval
- Better and Faster: Knowledge Transfer from Multiple Self-supervised Learning Tasks via Graph Distillation for Video Classification
- Fighting Fake News: Image Splice Detection via Learned Self-Consistency
- Joint Learning of Motion Estimation and Segmentation for Cardiac MR Image Sequences
- Learning Image Matching by Simply Watching Video
- Learning to Recognize Objects by Retaining other Factors of Variation
- Deep Lidar CNN to Understand the Dynamics of Moving Vehicles
- Actor and Observer: Joint Modeling of First and Third-Person Videos
- Customizing First Person Image Through Desired Actions
- Representation Learning by Reconstructing Neighborhoods
- ECO: Egocentric Cognitive Mapping
- Unsupervised Learning of Depth and Deep Representation for Visual Odometry from Monocular Videos in a Metric Space
- A Multi-view Perspective of Self-supervised Learning
- Learning View and Target Invariant Visual Servoing for Navigation
- Weakly-Supervised Spatial Context Networks
- Learning to Estimate Pose by Watching Videos
- A Classification approach towards Unsupervised Learning of Visual Representations
- Ambient Sound Provides Supervision for Visual Learning