Learning to Poke by Poking: Experiential Learning of Intuitive Physics
arXiv:1606.07419
Abstract
We investigate an experiential learning paradigm for acquiring an internal model of intuitive physics. Our model is evaluated on a real-world robotic manipulation task that requires displacing objects to target locations by poking. The robot gathered over 400 hours of experience by executing more than 100K pokes on different objects. We propose a novel approach based on deep neural networks for modeling the dynamics of robot's interactions directly from images, by jointly estimating forward and inverse models of dynamics. The inverse model objective provides supervision to construct informative visual features, which the forward model can then predict and in turn regularize the feature space for the inverse model. The interplay between these two objectives creates useful, accurate models that can then be used for multi-step decision making. This formulation has the additional benefit that it is possible to learn forward models in an abstract feature space and thus alleviate the need of predicting pixels. Our experiments show that this joint modeling approach outperforms alternative methods.
Cited by in corpus (55)
- Data-Efficient Image Recognition with Contrastive Predictive Coding
- Deep Reinforcement Learning: An Overview
- Visual Reinforcement Learning with Imagined Goals
- State Representation Learning for Control: An Overview
- Visual Foresight: Model-Based Deep Reinforcement Learning for Vision-Based Robotic Control
- Learning to Predict the Cosmological Structure Formation
- Flexible Neural Representation for Physics Prediction
- Self-Supervised Visual Planning with Temporal Skip Connections
- Loss is its own Reward: Self-Supervision for Reinforcement Learning
- Learning Predictive Representations for Deformable Objects Using Contrastive Estimation
- Universal Planning Networks
- A Survey of Deep Network Solutions for Learning Control in Robotics: From Reinforcement to Imitation
- Visual Interaction Networks
- Learning and Querying Fast Generative Models for Reinforcement Learning
- Contingency-Aware Exploration in Reinforcement Learning
- Learning to Perform Physics Experiments via Deep Reinforcement Learning
- Learning Dynamic Belief Graphs to Generalize on Text-Based Games
- Mid-Level Visual Representations Improve Generalization and Sample Efficiency for Learning Visuomotor Policies
- Sim2Real View Invariant Visual Servoing by Recurrent Control
- Temporal Relational Reasoning in Videos
- Decoupling Dynamics and Reward for Transfer Learning
- SE3-Pose-Nets: Structured Deep Dynamics Models for Visuomotor Planning and Control
- Learning to Fly by Crashing
- Learning Latent Plans from Play
- Deep Visual Foresight for Planning Robot Motion
- Prediction and Control with Temporal Segment Models
- Graph-Structured Visual Imitation
- Learning Robotic Manipulation through Visual Planning and Acting
- Robust Policies via Mid-Level Visual Representations: An Experimental Study in Manipulation and Navigation
- Forward-Backward Reinforcement Learning
- O2O-Afford: Annotation-Free Large-Scale Object-Object Affordance Learning
- 3D-OES: Viewpoint-Invariant Object-Factorized Environment Simulators
- Learning to Navigate Using Mid-Level Visual Priors
- Which Mutual-Information Representation Learning Objectives are Sufficient for Control?
- Learning Actionable Representations with Goal-Conditioned Policies
- A Long Horizon Planning Framework for Manipulating Rigid Pointcloud Objects
- Towards Lifelong Self-Supervision: A Deep Learning Direction for Robotics
- Video Jigsaw: Unsupervised Learning of Spatiotemporal Context for Video Action Recognition
- The "something something" video database for learning and evaluating visual common sense
- Low Dimensional State Representation Learning with Reward-shaped Priors
- Active Perception and Representation for Robotic Manipulation
- Learning Transferable Push Manipulation Skills in Novel Contexts
- Binge Watching: Scaling Affordance Learning from Sitcoms
- Acquiring Target Stacking Skills by Goal-Parameterized Deep Reinforcement Learning
- Physical Primitive Decomposition
- Interleaved Multitask Learning with Energy Modulated Learning Progress
- Towards Robust Bisimulation Metric Learning
- Instance-Aware Predictive Navigation in Multi-Agent Environments
- Swoosh! Rattle! Thump! -- Actions that Sound
- Invariant Feature Mappings for Generalizing Affordance Understanding Using Regularized Metric Learning
- Interpretable Intuitive Physics Model
- Bridging Cognitive Programs and Machine Learning
- Learning Rich Representations For Structured Visual Prediction Tasks
- Learning Sports Camera Selection from Internet Videos
- Hierarchical Neural Dynamic Policies