668 citations · 2k across the 48 of their papers we have counts for
46 papers · 1 filter
Learning from negative feedback, or positive feedback or both
Abbas Abdolmaleki, Bilal Piot, Bobak Shahriari +9
Existing preference optimization methods often assume scenarios where paired preference feedback (preferred/positive vs. dis-preferred/negative examples) is available. This require…
Replay across Experiments: A Natural Extension of Off-Policy RL
Dhruva Tirumala, Thomas Lampe, Jose Enrique Chen +9
Replaying data is a principal mechanism underlying the stability and data efficiency of off-policy reinforcement learning (RL). We present an effective yet simple framework to exte…
MO2: Model-Based Offline Options
Sasha Salter, Markus Wulfmeier, Dhruva Tirumala +4
The ability to discover useful behaviours from past experience and transfer them to new tasks is considered a core component of natural embodied intelligence. Inspired by neuroscie…
Data augmentation for efficient learning from parametric experts
Alexandre Galashov, Josh Merel, Nicolas Heess
We present a simple, yet powerful data-augmentation technique to enable data-efficient learning from parametric experts for reinforcement and imitation learning. We focus on what w…
Revisiting Gaussian mixture critics in off-policy reinforcement learning: a sample-based approach
Bobak Shahriari, Abbas Abdolmaleki, Arunkumar Byravan +6
Actor-critic algorithms that make use of distributional policy evaluation have frequently been shown to outperform their non-distributional counterparts on many challenging control…
COptiDICE: Offline Constrained Reinforcement Learning via Stationary Distribution Correction Estimation
Jongmin Lee, Cosmin Paduraru, Daniel J. Mankowitz +4
We consider the offline constrained reinforcement learning (RL) problem, in which the agent aims to compute a policy that maximizes expected return while satisfying given cost cons…