668 citations · 1.8k across the 15 of their papers we have counts for
4 papers · 1 filter
Learning Dynamics Models for Model Predictive Agents
Michael Lutter, Leonard Hasenclever, Arunkumar Byravan +5
Model-Based Reinforcement Learning involves learning a \textit{dynamics model} from data, and then using this model to optimise behaviour, most often with an online \textit{planner…
Local Search for Policy Iteration in Continuous Control
Jost Tobias Springenberg, Nicolas Heess, Daniel Mankowitz +10
We present an algorithm for local, regularized, policy improvement in reinforcement learning (RL) that allows us to formulate model-based and model-free variants in a single framew…
Relative Entropy Regularized Policy Iteration
Abbas Abdolmaleki, Jost Tobias Springenberg, Jonas Degrave +5
We present an off-policy actor-critic algorithm for Reinforcement Learning (RL) that combines ideas from gradient-free optimization via stochastic search with learned action-value…
Maximum a Posteriori Policy Optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa +3
We introduce a new algorithm for reinforcement learning called Maximum aposteriori Policy Optimisation (MPO) based on coordinate ascent on a relative entropy objective. We show tha…