Unsupervised Learning for Physical Interaction through Video Prediction
arXiv:1605.07157
Abstract
A core challenge for an agent learning to interact with the world is to predict how its actions affect objects in its environment. Many existing methods for learning the dynamics of physical interactions require labeled object information. However, to scale real-world interaction learning to a variety of scenes and objects, acquiring labeled data becomes increasingly impractical. To learn about physical object motion without labels, we develop an action-conditioned video prediction model that explicitly models pixel motion, by predicting a distribution over pixel motion from previous frames. Because our model explicitly predicts motion, it is partially invariant to object appearance, enabling it to generalize to previously unseen objects. To explore video prediction for real-world interactive agents, we also introduce a dataset of 59,000 robot interactions involving pushing motions, including a test set with novel objects. In this dataset, accurate prediction of videos conditioned on the robot's future actions amounts to learning a "visual imagination" of different futures based on different courses of action. Our experiments show that our proposed method produces more accurate video predictions both quantitatively and qualitatively, when compared to prior methods.
To appear in NIPS '16; Video results, code, and data available at: http://www.sites.google.com/site/robotprediction
References in corpus (3)
Cited by in corpus (87)
- Deep Learning for Precipitation Nowcasting: A Benchmark and A New Model
- Decomposing Motion and Content for Natural Video Sequence Prediction
- Deep Predictive Coding Networks for Video Prediction and Unsupervised Learning
- State Representation Learning for Control: An Overview
- Deep Learning for Physical Processes: Incorporating Prior Scientific Knowledge
- Visual Foresight: Model-Based Deep Reinforcement Learning for Vision-Based Robotic Control
- Stochastic Adversarial Video Prediction
- Flexible Neural Representation for Physics Prediction
- Visual Dynamics: Probabilistic Future Frame Synthesis via Cross Convolutional Networks
- Video-to-Video Synthesis
- A Disentangled Recognition and Nonlinear Dynamics Model for Unsupervised Learning
- Train Sparsely, Generate Densely: Memory-efficient Unsupervised Training of High-resolution Temporal GAN
- Self-Supervised Visual Planning with Temporal Skip Connections
- Learning to Decompose and Disentangle Representations for Video Prediction
- Photographic Image Synthesis with Cascaded Refinement Networks
- GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions
- Stochastic Variational Video Prediction
- CM-GANs: Cross-modal Generative Adversarial Networks for Common Representation Learning
- Transformation-Based Models of Video Sequences
- Generative OpenMax for Multi-Class Open Set Classification
- Learning Video Object Segmentation with Visual Memory
- Unsupervised learning of foreground object detection
- DeepPhys: Video-Based Physiological Measurement Using Convolutional Attention Networks
- Value Prediction Network
- Predictive-Corrective Networks for Action Detection
- End-to-end Learning of Driving Models from Large-scale Video Datasets
- Transformation-Grounded Image Generation Network for Novel 3D View Synthesis
- Predicting Deeper into the Future of Semantic Segmentation
- Detecting Adversarial Attacks on Neural Network Policies with Visual Foresight
- Conditional Adversarial Network for Semantic Segmentation of Brain Tumor
- Future Semantic Segmentation with Convolutional LSTM
- Object Detection in Video with Spatiotemporal Sampling Networks
- Accurate and Diverse Sampling of Sequences based on a "Best of Many" Sample Objective
- Time-Agnostic Prediction: Predicting Predictable Video Frames
- How do Mixture Density RNNs Predict the Future?
- Deep Visual Foresight for Planning Robot Motion
- Approximating the solution to wave propagation using deep neural networks
- Occlusion Aware Unsupervised Learning of Optical Flow
- Dynamics Transfer GAN: Generating Video by Transferring Arbitrary Temporal Dynamics from a Source Video to a Single Target Image
- Prediction Under Uncertainty with Error-Encoding Networks
- Unsupervised Learning from Continuous Video in a Scalable Predictive Recurrent Network
- Hybrid Learning of Optical Flow and Next Frame Prediction to Boost Optical Flow in the Wild
- Novel Video Prediction for Large-scale Scene using Optical Flow
- Improving Video Generation for Multi-functional Applications
- Reduced-Gate Convolutional LSTM Using Predictive Coding for Spatiotemporal Prediction
- Flow-Grounded Spatial-Temporal Video Prediction from Still Images
- Im2Flow: Motion Hallucination from Static Images for Action Recognition
- Unsupervised Feature Learning for Audio Analysis
- FASTER Recurrent Networks for Efficient Video Classification
- 3D-PhysNet: Learning the Intuitive Physics of Non-Rigid Object Deformations
- Super-Resolution with Deep Adaptive Image Resampling
- Novel View Synthesis for Large-scale Scene using Adversarial Loss
- Learning 6-DOF Grasping Interaction via Deep Geometry-aware 3D Representations
- Few-shot Video-to-Video Synthesis
- Relational Action Forecasting
- Visual Forecasting by Imitating Dynamics in Natural Sequences
- Adversarial Video Compression Guided by Soft Edge Detection
- Sequential Learning of Movement Prediction in Dynamic Environments using LSTM Autoencoder
- Independent Innovation Analysis for Nonlinear Vector Autoregressive Process
- Learning Robot Activities from First-Person Human Videos Using Convolutional Future Regression
- Attentioned Convolutional LSTM InpaintingNetwork for Anomaly Detection in Videos
- On Machine Learning and Structure for Mobile Robots
- What Would You Do? Acting by Learning to Predict
- SCH-GAN: Semi-supervised Cross-modal Hashing by Generative Adversarial Network
- DRL: Deep Reinforcement Learning for Intelligent Robot Control -- Concept, Literature, and Future
- Adversarial Framework for Unsupervised Learning of Motion Dynamics in Videos
- ContextVP: Fully Context-Aware Video Prediction
- Image2GIF: Generating Cinemagraphs using Recurrent Deep Q-Networks
- Long-Term Image Boundary Prediction
- Seeing in the dark with recurrent convolutional neural networks
- Personalized Saliency and its Prediction
- Learning to Forecast Videos of Human Activity with Multi-granularity Models and Adaptive Rendering
- Human Pose Forecasting via Deep Markov Models
- Learning Temporal Dynamics from Cycles in Narrated Video
- Recurrent Multimodal Interaction for Referring Image Segmentation
- Motion Selective Prediction for Video Frame Synthesis
- Recurrent Flow-Guided Semantic Forecasting
- Modular meta-learning in abstract graph networks for combinatorial generalization
- Forecasting Hands and Objects in Future Frames
- Interpretable Intuitive Physics Model
- Inserting Videos into Videos
- Predicting the Future with Transformational States
- Defo-Net: Learning Body Deformation using Generative Adversarial Networks
- MT-VAE: Learning Motion Transformations to Generate Multimodal Human Dynamics
- Transferring Agent Behaviors from Videos via Motion GANs
- VP-GO: a "light" action-conditioned visual prediction model
- Neural Embedding for Physical Manipulations