Time-Contrastive Networks: Self-Supervised Learning from Video
arXiv:1704.06888
Abstract
We propose a self-supervised approach for learning representations and robotic behaviors entirely from unlabeled videos recorded from multiple viewpoints, and study how this representation can be used in two robotic imitation settings: imitating object interactions from videos of humans, and imitating human poses. Imitation of human behavior requires a viewpoint-invariant representation that captures the relationships between end-effectors (hands or robot grippers) and the environment, object attributes, and body pose. We train our representations using a metric learning loss, where multiple simultaneous viewpoints of the same observation are attracted in the embedding space, while being repelled from temporal neighbors which are often visually similar but functionally different. In other words, the model simultaneously learns to recognize what is common between different-looking images, and what is different between similar-looking images. This signal causes our model to discover attributes that do not change across viewpoint, but do change across time, while ignoring nuisance variables such as occlusions, motion blur, lighting and background. We demonstrate that this representation can be used by a robot to directly mimic human poses without an explicit correspondence, and that it can be used as a reward function within a reinforcement learning algorithm. While representations are learned from an unlabeled collection of task-related videos, robot behaviors such as pouring are learned by watching a single 3rd-person demonstration by a human. Reward functions obtained by following the human demonstrations under the learned representation enable efficient reinforcement learning that is practical for real-world robotic systems. Video results, open-source code and dataset are available at https://sermanet.github.io/imitate
References in corpus (16)
- Adversarially Learned Inference
- Rethinking the Inception Architecture for Computer Vision
- Unsupervised Visual Representation Learning by Context Prediction
- SoundNet: Learning Sound Representations from Unlabeled Video
- One-Shot Imitation Learning
- Learning to Compare Image Patches via Convolutional Neural Networks
- Unsupervised Learning of Visual Representations using Videos
- Third-Person Imitation Learning
- LIFT: Learned Invariant Feature Transform
- Generic 3D Representation via Pose Estimation and Matching
- Imitation from Observation: Learning to Imitate Behaviors from Raw Video via Context Translation
- Understanding Visual Concepts with Continuation Learning
- Pose Embeddings: A Deep Architecture for Learning to Match Human Poses
- Learning Features by Watching Objects Move
- Split-Brain Autoencoders: Unsupervised Learning by Cross-Channel Prediction
- Unsupervised Perceptual Rewards for Imitation Learning
Cited by in corpus (33)
- Data-Efficient Image Recognition with Contrastive Predictive Coding
- Learning Representations by Maximizing Mutual Information Across Views
- Visual Reinforcement Learning with Imagined Goals
- Episodic Curiosity through Reachability
- One-Shot Imitation from Observing Humans via Domain-Adaptive Meta-Learning
- Reinforcement and Imitation Learning for Diverse Visuomotor Skills
- Meta-Learning Update Rules for Unsupervised Representation Learning
- Universal Planning Networks
- Revisiting Self-Supervised Visual Representation Learning
- Imitating Latent Policies from Observation
- Imitation from Observation: Learning to Imitate Behaviors from Raw Video via Context Translation
- Temporal Relational Reasoning in Videos
- Robot eye-hand coordination learning by watching human demonstrations: a task function approximation approach
- S3K: Self-Supervised Semantic Keypoints for Robotic Manipulation via Multi-View Consistency
- Neural Task Graphs: Generalizing to Unseen Tasks from a Single Video Demonstration
- Multigrid Predictive Filter Flow for Unsupervised Learning on Videos
- Mine Your Own vieW: Self-Supervised Learning Through Across-Sample Prediction
- Learning Object Manipulation Skills via Approximate State Estimation from Real Videos
- Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions
- XIRL: Cross-embodiment Inverse Reinforcement Learning
- Self-supervisory Signals for Object Discovery and Detection
- Vision-based Teleoperation of Shadow Dexterous Hand using End-to-End Deep Neural Network
- Temporal Difference Learning with Neural Networks - Study of the Leakage Propagation Problem
- Estimating Q(s,s') with Deep Deterministic Dynamics Gradients
- StackMix: A complementary Mix algorithm
- Learning to Align Sequential Actions in the Wild
- Reparameterized Variational Divergence Minimization for Stable Imitation
- Towards Object Detection from Motion
- State Representation Learning from Demonstration
- Transferring Agent Behaviors from Videos via Motion GANs
- Supervise Thyself: Examining Self-Supervised Representations in Interactive Environments
- MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics
- Injective State-Image Mapping facilitates Visual Adversarial Imitation Learning