Path Integral Networks: End-to-End Differentiable Optimal Control
arXiv:1706.09597
Abstract
In this paper, we introduce Path Integral Networks (PI-Net), a recurrent network representation of the Path Integral optimal control algorithm. The network includes both system dynamics and cost models, used for optimal control based planning. PI-Net is fully differentiable, learning both dynamics and cost models end-to-end by back-propagation and stochastic gradient descent. Because of this, PI-Net can learn to plan. PI-Net has several advantages: it can generalize to unseen states thanks to planning, it can be applied to continuous control tasks, and it allows for a wide variety learning schemes, including imitation and reinforcement learning. Preliminary experiment results show that PI-Net, trained by imitation learning, can mimic control demonstrations for two simulated problems; a linear system and a pendulum swing-up problem. We also show that PI-Net is able to learn dynamics and cost models latent in the demonstrations.
Cited by in corpus (16)
- Differentiable MPC for End-to-end Planning and Control
- Learning to Guide: Guidance Law Based on Deep Meta-learning and Model Predictive Path Integral Control
- QMDP-Net: Deep Learning for Planning under Partial Observability
- Objective Mismatch in Model-based Reinforcement Learning
- Acceleration of Gradient-based Path Integral Method for Efficient Optimal and Inverse Optimal Control
- Particle Filter Networks with Application to Visual Localization
- On the model-based stochastic value gradient for continuous reinforcement learning
- Learning to Optimize in Model Predictive Control
- Pontryagin Differentiable Programming: An End-to-End Learning and Control Framework
- PlaNet of the Bayesians: Reconsidering and Improving Deep Planning Network by Incorporating Bayesian Inference
- Learning Convex Optimization Control Policies
- Disentangling Controllable and Uncontrollable Factors of Variation by Interacting with the World
- The Differentiable Cross-Entropy Method
- Integrating Algorithmic Planning and Deep Learning for Partially Observable Navigation
- Learning Navigation Costs from Demonstration in Partially Observable Environments
- Distilling a Hierarchical Policy for Planning and Control via Representation and Reinforcement Learning