activity
20162022
most citedEmergence of Locomotion Behaviours in Rich Environments

668 citations · 1.5k across the 12 of their papers we have counts for

collaborators
Showing cs.LGShow all

8 papers · 1 filter

cs.LG20222 cited

Data augmentation for efficient learning from parametric experts

Alexandre Galashov, Josh Merel, Nicolas Heess

We present a simple, yet powerful data-augmentation technique to enable data-efficient learning from parametric experts for reinforcement and imitation learning. We focus on what w…

cs.LG20214 cited

Learning Dynamics Models for Model Predictive Agents

Michael Lutter, Leonard Hasenclever, Arunkumar Byravan +5

Model-Based Reinforcement Learning involves learning a \textit{dynamics model} from data, and then using this model to optimise behaviour, most often with an online \textit{planner…

cs.LG2020

Local Search for Policy Iteration in Continuous Control

Jost Tobias Springenberg, Nicolas Heess, Daniel Mankowitz +10

We present an algorithm for local, regularized, policy improvement in reinforcement learning (RL) that allows us to formulate model-based and model-free variants in a single framew…

cs.LG2020

RL Unplugged: A Suite of Benchmarks for Offline Reinforcement Learning

Caglar Gulcehre, Ziyu Wang, Alexander Novikov +15

Offline methods for reinforcement learning have a potential to help bridge the gap between reinforcement learning research and real-world applications. They make it possible to lea…

cs.LG20209 cited

Divide-and-Conquer Monte Carlo Tree Search For Goal-Directed Planning

Giambattista Parascandolo, Lars Buesing, Josh Merel +6

Standard planners for sequential decision making (including Monte Carlo planning, tree search, dynamic programming, etc.) are constrained by an implicit sequential planning assumpt…

cs.LG2018

Neural probabilistic motor primitives for humanoid control

Josh Merel, Leonard Hasenclever, Alexandre Galashov +5

We focus on the problem of learning a single motor module that can flexibly express a range of behaviors for the control of high-dimensional physically simulated humanoids. To do t…