On Effective Scheduling of Model-based Reinforcement Learning
arXiv:2111.08550
Abstract
Model-based reinforcement learning has attracted wide attention due to its superior sample efficiency. Despite its impressive success so far, it is still unclear how to appropriately schedule the important hyperparameters to achieve adequate performance, such as the real data ratio for policy optimization in Dyna-style model-based algorithms. In this paper, we first theoretically analyze the role of real data in policy training, which suggests that gradually increasing the ratio of real data yields better performance. Inspired by the analysis, we propose a framework named AutoMBPO to automatically schedule the real data ratio as well as other hyperparameters in training model-based policy optimization (MBPO) algorithm, a representative running case of model-based methods. On several continuous control tasks, the MBPO instance trained with hyperparameters scheduled by AutoMBPO can significantly surpass the original one, and the real data ratio schedule found by AutoMBPO shows consistency with our theoretical analysis.
Accepted at NeurIPS2021
References in corpus (15)
- Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor
- Model-Ensemble Trust-Region Policy Optimization
- Benchmarking Model-Based Reinforcement Learning
- Model-Based Value Estimation for Efficient Model-Free Reinforcement Learning
- EPOpt: Learning Robust Neural Network Policies Using Model Ensembles
- Algorithmic Framework for Model-based Deep Reinforcement Learning with Theoretical Guarantees
- Model-Based Reinforcement Learning via Meta-Policy Optimization
- Model-based RL in Contextual Decision Processes: PAC bounds and Exponential Improvements over Model-free Approaches
- A Game Theoretic Framework for Model Based Reinforcement Learning
- Model-Augmented Actor-Critic: Backpropagating through Paths
- On the Importance of Hyperparameter Optimization for Model-based Reinforcement Learning
- AutoLoss: Learning Discrete Schedules for Alternate Optimization
- Trust the Model When It Is Confident: Masked Model-based Actor-Critic
- Bidirectional Model-based Policy Optimization
- Bridging Imagination and Reality for Model-Based Deep Reinforcement Learning