Data-Efficient Reinforcement Learning with Probabilistic Model Predictive Control
arXiv:1706.06491
Abstract
Trial-and-error based reinforcement learning (RL) has seen rapid advancements in recent times, especially with the advent of deep neural networks. However, the majority of autonomous RL algorithms require a large number of interactions with the environment. A large number of interactions may be impractical in many real-world applications, such as robotics, and many practical systems have to obey limitations in the form of state space or control constraints. To reduce the number of system interactions while simultaneously handling constraints, we propose a model-based RL framework based on probabilistic Model Predictive Control (MPC). In particular, we propose to learn a probabilistic transition model using Gaussian Processes (GPs) to incorporate model uncertainty into long-term predictions, thereby, reducing the impact of model errors. We then use MPC to find a control sequence that minimises the expected long-term cost. We provide theoretical guarantees for first-order optimality in the GP-based transition models with deterministic approximate inference for long-term planning. We demonstrate that our approach does not only achieve state-of-the-art data efficiency, but also is a principled way for RL in constrained environments.
Accepted at AISTATS 2018,
References in corpus (2)
Cited by in corpus (33)
- Comparison of Deep Reinforcement Learning and Model Predictive Control for Adaptive Cruise Control
- Meta Reinforcement Learning with Latent Variable Gaussian Processes
- Model-Based Meta-Reinforcement Learning for Flight with Suspended Payloads
- Dynamics-Aware Unsupervised Discovery of Skills
- Generalizing Convolutional Neural Networks for Equivariance to Lie Groups on Arbitrary Continuous Data
- Learning-based Model Predictive Control for Safe Exploration and Reinforcement Learning
- Learning to Guide: Guidance Law Based on Deep Meta-learning and Model Predictive Path Integral Control
- On Simulation and Trajectory Prediction with Gaussian Process Dynamics
- Convergence results for an averaged LQR problem with applications to reinforcement learning
- Efficiently Sampling Functions from Gaussian Process Posteriors
- A Tutorial on Sparse Gaussian Processes and Variational Inference
- Application of the Free Energy Principle to Estimation and Control
- Model-based Lookahead Reinforcement Learning
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and Planning
- Deep Model-Based Reinforcement Learning for High-Dimensional Problems, a Survey
- Reinforcement Learning Ship Autopilot: Sample efficient and Model Predictive Control-based Approach
- Learning ODE Models with Qualitative Structure Using Gaussian Processes
- Model Imitation for Model-Based Reinforcement Learning
- Continuous-Time Model-Based Reinforcement Learning
- Structured Variational Inference in Unstable Gaussian Process State Space Models
- Pathwise Conditioning of Gaussian Processes
- Benchmarking Structured Policies and Policy Optimization for Real-World Dexterous Object Manipulation
- Combining Reinforcement Learning with Model Predictive Control for On-Ramp Merging
- Meta Learning MPC using Finite-Dimensional Gaussian Process Approximations
- VMAV-C: A Deep Attention-based Reinforcement Learning Algorithm for Model-based Control
- Safe Policy Search with Gaussian Process Models
- Learning to Plan Optimistically: Uncertainty-Guided Deep Exploration via Latent Model Ensembles
- Model-based Multi-agent Policy Optimization with Adaptive Opponent-wise Rollouts
- Uncertainty-aware Contact-safe Model-based Reinforcement Learning
- Model-based Meta Reinforcement Learning using Graph Structured Surrogate Models
- PAC Bounds for Imitation and Model-based Batch Learning of Contextual Markov Decision Processes
- Data-Efficient Reinforcement Learning for Malaria Control
- Distributionally Robust Trajectory Optimization Under Uncertain Dynamics via Relative Entropy Trust-Regions