668 citations · 2.5k across the 82 of their papers we have counts for
10 papers · 2 filters
Relative Entropy Regularized Policy Iteration
Abbas Abdolmaleki, Jost Tobias Springenberg, Jonas Degrave +5
We present an off-policy actor-critic algorithm for Reinforcement Learning (RL) that combines ideas from gradient-free optimization via stochastic search with learned action-value…
Composing Entropic Policies using Divergence Correction
Jonathan J Hunt, Andre Barreto, Timothy P Lillicrap +1
Composing previously mastered skills to solve novel tasks promises dramatic improvements in the data efficiency of reinforcement learning. Here, we analyze two recent works composi…
Woulda, Coulda, Shoulda: Counterfactually-Guided Policy Search
Lars Buesing, Theophane Weber, Yori Zwols +4
Learning policies on data synthesized by models can in principle quench the thirst of reinforcement learning algorithms for large amounts of real experience, which is often costly…
Neural probabilistic motor primitives for humanoid control
Josh Merel, Leonard Hasenclever, Alexandre Galashov +5
We focus on the problem of learning a single motor module that can flexibly express a range of behaviors for the control of high-dimensional physically simulated humanoids. To do t…
Maximum a Posteriori Policy Optimisation
Abbas Abdolmaleki, Jost Tobias Springenberg, Yuval Tassa +3
We introduce a new algorithm for reinforcement learning called Maximum aposteriori Policy Optimisation (MPO) based on coordinate ascent on a relative entropy objective. We show tha…
Mix&Match - Agent Curricula for Reinforcement Learning
Wojciech Marian Czarnecki, Siddhant M. Jayakumar, Max Jaderberg +5
We introduce Mix&Match (M&M) - a training framework designed to facilitate rapid and effective learning in RL agents, especially those that would be too slow or too challenging to…