Gradient-Aware Model-based Policy Search
arXiv:1909.04115 · doi:10.1609/aaai.v34i04.5791
Abstract
Traditional model-based reinforcement learning approaches learn a model of the environment dynamics without explicitly considering how it will be used by the agent. In the presence of misspecified model classes, this can lead to poor estimates, as some relevant available information is ignored. In this paper, we introduce a novel model-based policy search approach that exploits the knowledge of the current agent policy to learn an approximate transition model, focusing on the portions of the environment that are most relevant for policy improvement. We leverage a weighting scheme, derived from the minimization of the error on the model-based policy gradient estimator, in order to define a suitable objective function that is optimized for learning the approximate transition model. Then, we integrate this procedure into a batch policy improvement algorithm, named Gradient-Aware Model-based Policy Search (GAMPS), which iteratively learns a transition model and uses it, together with the collected trajectories, to compute the new policy parameters. Finally, we empirically validate GAMPS on benchmark domains analyzing and discussing its properties.
References in corpus (14)
- Adam: A Method for Stochastic Optimization
- Recurrent World Models Facilitate Policy Evolution
- Learning Continuous Control Policies by Stochastic Value Gradients
- Model-Ensemble Trust-Region Policy Optimization
- Differentiable MPC for End-to-end Planning and Control
- Deep Reinforcement Learning in a Handful of Trials using Probabilistic Dynamics Models
- Model-Based Value Estimation for Efficient Model-Free Reinforcement Learning
- The Predictron: End-To-End Learning and Planning
- Sample-Efficient Reinforcement Learning with Stochastic Ensemble Value Expansion
- Algorithmic Framework for Model-based Deep Reinforcement Learning with Theoretical Guarantees
- Task-based End-to-end Model Learning in Stochastic Optimization
- Agnostic System Identification for Model-Based Reinforcement Learning
- Goal-Driven Dynamics Learning via Bayesian Optimization
- Equivalence Between Wasserstein and Value-Aware Loss for Model-based Reinforcement Learning