When to Trust Your Model: Model-Based Policy Optimization
arXiv:1906.08253
Abstract
Designing effective model-based reinforcement learning algorithms is difficult because the ease of data generation must be weighed against the bias of model-generated data. In this paper, we study the role of model usage in policy optimization both theoretically and empirically. We first formulate and analyze a model-based reinforcement learning algorithm with a guarantee of monotonic improvement at each step. In practice, this analysis is overly pessimistic and suggests that real off-policy data is always preferable to model-generated on-policy data, but we show that an empirical estimate of model generalization can be incorporated into such analysis to justify model usage. Motivated by this analysis, we then demonstrate that a simple procedure of using short model-generated rollouts branched from real data has the benefits of more complicated model-based algorithms without the usual pitfalls. In particular, this approach surpasses the sample efficiency of prior model-based methods, matches the asymptotic performance of the best model-free algorithms, and scales to horizons that cause other model-based methods to fail entirely.
NeurIPS 2019. Code at https://github.com/JannerM/mbpo, project page at: https://jannerm.github.io/mbpo-www/
References in corpus (11)
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Model-Based Reinforcement Learning for Atari
- Model-Ensemble Trust-Region Policy Optimization
- Visual Foresight: Model-Based Deep Reinforcement Learning for Vision-Based Robotic Control
- Model-Based Value Estimation for Efficient Model-Free Reinforcement Learning
- Sample-Efficient Reinforcement Learning with Stochastic Ensemble Value Expansion
- Algorithmic Framework for Model-based Deep Reinforcement Learning with Theoretical Guarantees
- Model-Based Reinforcement Learning via Meta-Policy Optimization
- The Effect of Planning Shape on Dyna-style Planning in High-dimensional State Spaces
- Dual Policy Iteration
- Task-Agnostic Dynamics Priors for Deep Reinforcement Learning
Cited by in corpus (30)
- MOReL : Model-Based Offline Reinforcement Learning
- What Matters In On-Policy Reinforcement Learning? A Large-Scale Empirical Study
- Learning to Walk in the Real World with Minimal Human Effort
- Automated Reinforcement Learning (AutoRL): A Survey and Open Problems
- A Game Theoretic Framework for Model Based Reinforcement Learning
- Model-Augmented Actor-Critic: Backpropagating through Paths
- Error Bounds of Imitating Policies and Environments
- Dexterous Robotic Manipulation using Deep Reinforcement Learning and Knowledge Transfer for Complex Sparse Reward-based Tasks
- On the model-based stochastic value gradient for continuous reinforcement learning
- Adaptive Online Planning for Continual Lifelong Learning
- Efficient Model-Based Reinforcement Learning through Optimistic Policy Search and Planning
- Ready Policy One: World Building Through Active Learning
- Trajectory-wise Multiple Choice Learning for Dynamics Generalization in Reinforcement Learning
- MUSBO: Model-based Uncertainty Regularized and Sample Efficient Batch Optimization for Deployment Constrained Reinforcement Learning
- Model-free and Bayesian Ensembling Model-based Deep Reinforcement Learning for Particle Accelerator Control Demonstrated on the FERMI FEL
- MobILE: Model-Based Imitation Learning From Observation Alone
- Mean-Semivariance Policy Optimization via Risk-Averse Reinforcement Learning
- Approximate Model-Based Shielding for Safe Reinforcement Learning
- Building Minimal and Reusable Causal State Abstractions for Reinforcement Learning
- CostNet: An End-to-End Framework for Goal-Directed Reinforcement Learning
- Maximum Entropy Model Rollouts: Fast Model Based Policy Optimization without Compounding Errors
- Monte Carlo Tree Search in the Presence of Transition Uncertainty
- Learning Off-Policy with Online Planning
- Doubly Robust Off-Policy Actor-Critic Algorithms for Reinforcement Learning
- GEM: Group Enhanced Model for Learning Dynamical Control Systems
- Placement Optimization with Deep Reinforcement Learning
- Reward Balancing Revisited: Enhancing Offline Reinforcement Learning for Recommender Systems
- DROMO: Distributionally Robust Offline Model-based Policy Optimization
- On-Policy Model Errors in Reinforcement Learning
- Minimax Model Learning