Prediction and Control with Temporal Segment Models
arXiv:1703.04070
Abstract
We introduce a method for learning the dynamics of complex nonlinear systems based on deep generative models over temporal segments of states and actions. Unlike dynamics models that operate over individual discrete timesteps, we learn the distribution over future state trajectories conditioned on past state, past action, and planned future action trajectories, as well as a latent prior over action trajectories. Our approach is based on convolutional autoregressive models and variational autoencoders. It makes stable and accurate predictions over long horizons for complex, stochastic systems, effectively expressing uncertainty and modeling the effects of collisions, sensory noise, and action delays. The learned dynamics model and action prior can be used for end-to-end, fully differentiable trajectory optimization and model-based policy optimization, which we use to evaluate the performance and sample-efficiency of our method.
camera-ready version, ICML 2017
References in corpus (7)
- WaveNet: A Generative Model for Raw Audio
- Conditional Image Generation with PixelCNN Decoders
- Scheduled Sampling for Sequence Prediction with Recurrent Neural Networks
- Composing graphical models with neural networks for structured representations and fast inference
- Learning to Poke by Poking: Experiential Learning of Intuitive Physics
- Learning Visual Predictive Models of Physics for Playing Billiards
- Deep Visual Foresight for Planning Robot Motion
Cited by in corpus (10)
- Model-Ensemble Trust-Region Policy Optimization
- Self-Consistent Trajectory Autoencoder: Hierarchical Reinforcement Learning with Trajectory Embeddings
- Value Prediction Network
- MPC-Inspired Neural Network Policies for Sequential Decision Making
- Learning Dynamics Model in Reinforcement Learning by Incorporating the Long Term Future
- PLAS: Latent Action Space for Offline Reinforcement Learning
- Combining Model-Based and Model-Free Methods for Nonlinear Control: A Provably Convergent Policy Gradient Approach
- MUSBO: Model-based Uncertainty Regularized and Sample Efficient Batch Optimization for Deployment Constrained Reinforcement Learning
- Planning with Exploration: Addressing Dynamics Bottleneck in Model-based Reinforcement Learning
- Episodic Memory for Learning Subjective-Timescale Models