activity
20112025
most citedEmergence of Locomotion Behaviours in Rich Environments

668 citations · 2k across the 48 of their papers we have counts for

collaborators
Showing cs.LGShow all

46 papers · 1 filter

cs.LG2024

Learning from negative feedback, or positive feedback or both

Abbas Abdolmaleki, Bilal Piot, Bobak Shahriari +9

Existing preference optimization methods often assume scenarios where paired preference feedback (preferred/positive vs. dis-preferred/negative examples) is available. This require…

cs.LG20231 cited

Replay across Experiments: A Natural Extension of Off-Policy RL

Dhruva Tirumala, Thomas Lampe, Jose Enrique Chen +9

Replaying data is a principal mechanism underlying the stability and data efficiency of off-policy reinforcement learning (RL). We present an effective yet simple framework to exte…

cs.LG20221 cited

MO2: Model-Based Offline Options

Sasha Salter, Markus Wulfmeier, Dhruva Tirumala +4

The ability to discover useful behaviours from past experience and transfer them to new tasks is considered a core component of natural embodied intelligence. Inspired by neuroscie…

cs.LG20222 cited

Data augmentation for efficient learning from parametric experts

Alexandre Galashov, Josh Merel, Nicolas Heess

We present a simple, yet powerful data-augmentation technique to enable data-efficient learning from parametric experts for reinforcement and imitation learning. We focus on what w…

cs.LG2022

Revisiting Gaussian mixture critics in off-policy reinforcement learning: a sample-based approach

Bobak Shahriari, Abbas Abdolmaleki, Arunkumar Byravan +6

Actor-critic algorithms that make use of distributional policy evaluation have frequently been shown to outperform their non-distributional counterparts on many challenging control…

cs.LG20223 cited

COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation

Jongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz +4

We consider the offline constrained reinforcement learning (RL) problem, in which the agent aims to compute a policy that maximizes expected return while satisfying given cost cons…