MBMF: Model-Based Priors for Model-Free Reinforcement Learning
arXiv:1709.03153
Abstract
Reinforcement Learning is divided in two main paradigms: model-free and model-based. Each of these two paradigms has strengths and limitations, and has been successfully applied to real world domains that are appropriate to its corresponding strengths. In this paper, we present a new approach aimed at bridging the gap between these two paradigms. We aim to take the best of the two paradigms and combine them in an approach that is at the same time data-efficient and cost-savvy. We do so by learning a probabilistic dynamics model and leveraging it as a prior for the intertwined model-free optimization. As a result, our approach can exploit the generality and structure of the dynamics model, but is also capable of ignoring its inevitable inaccuracies, by directly incorporating the evidence provided by the direct observation of the cost. Preliminary results demonstrate that our approach outperforms purely model-based and model-free approaches, as well as the approach of simply switching from a model-based to a model-free setting.
After we submitted the paper for consideration in CoRL 2017 we found a paper published in the recent past with a similar method (see related work for a discussion). Considering the similarities between the two papers, we have decided to retract our paper from CoRL 2017
References in corpus (3)
Cited by in corpus (9)
- Differentiable MPC for End-to-end Planning and Control
- Combining Model-Based and Model-Free Methods for Nonlinear Control: A Provably Convergent Policy Gradient Approach
- Sample-Efficient Learning of Nonprehensile Manipulation Policies via Physics-Based Informed State Distributions
- Thief, Beware of What Get You There: Towards Understanding Model Extraction Attack
- Dyna-AIL : Adversarial Imitation Learning by Planning
- Combining Model-Free Q-Ensembles and Model-Based Approaches for Informed Exploration
- Accelerating Goal-Directed Reinforcement Learning by Model Characterization
- Mixed Reinforcement Learning with Additive Stochastic Uncertainty
- Guiding Robot Exploration in Reinforcement Learning via Automated Planning