activity
20172022
most citedEmergence of Locomotion Behaviours in Rich Environments

668 citations · 1.4k across the 21 of their papers we have counts for

collaborators

31 papers

cs.LG20221 cited

MO2: Model-Based Offline Options

Sasha Salter, Markus Wulfmeier, Dhruva Tirumala +4

The ability to discover useful behaviours from past experience and transfer them to new tasks is considered a core component of natural embodied intelligence. Inspired by neuroscie…

cs.LG2022

Revisiting Gaussian mixture critics in off-policy reinforcement learning: a sample-based approach

Bobak Shahriari, Abbas Abdolmaleki, Arunkumar Byravan +6

Actor-critic algorithms that make use of distributional policy evaluation have frequently been shown to outperform their non-distributional counterparts on many challenging control…

cs.LG20228 cited

The Challenges of Exploration for Offline Reinforcement Learning

Nathan Lambert, Markus Wulfmeier, William Whitney +5

Offline Reinforcement Learning (ORL) enablesus to separately study the two interlinked processes of reinforcement learning: collecting informative experience and inferring optimal…

cs.LG20213 cited

Is Bang-Bang Control All You Need? Solving Continuous Control with Bernoulli Policies

Tim Seyde, Igor Gilitschenski, Wilko Schwarting +4

Reinforcement learning (RL) for continuous control typically employs distributions whose support covers the entire action space. In this work, we investigate the colloquially known…

cs.RO202116 cited

Beyond Pick-and-Place: Tackling Robotic Stacking of Diverse Shapes

Alex X. Lee, Coline Devin, Yuxiang Zhou +18

We study the problem of robotic stacking with objects of complex geometry. We propose a challenging and diverse set of such objects that was carefully designed to require strategie…

cs.RO20211 cited

Evaluating model-based planning and planner amortization for continuous control

Arunkumar Byravan, Leonard Hasenclever, Piotr Trochim +8

There is a widespread intuition that model-based control methods should be able to surpass the data efficiency of model-free approaches. In this paper we attempt to evaluate this i…