activity
20112025
most citedEmergence of Locomotion Behaviours in Rich Environments

668 citations · 2k across the 61 of their papers we have counts for

collaborators
Showing 2019Show all

18 papers · 1 filter

cs.LG201917 cited

Hindsight Credit Assignment

Anna Harutyunyan, Will Dabney, Thomas Mesnard +8

We consider the problem of efficient credit assignment in reinforcement learning. In order to efficiently and meaningfully utilize new data, we propose to explicitly assign credit…

cs.LG20194 cited

Quinoa: a Q-function You Infer Normalized Over Actions

Jonas Degrave, Abbas Abdolmaleki, Jost Tobias Springenberg +2

We present an algorithm for learning an approximate action-value soft Q-function in the relative entropy regularised reinforcement learning setting, for which an optimal improved p…

cs.AI2019

Catch & Carry: Reusable Neural Controllers for Vision-Guided Whole-Body Tasks

Josh Merel, Saran Tunyasuvunakool, Arun Ahuja +6

We address the longstanding challenge of producing flexible, realistic humanoid character controllers that can perform diverse whole-body tasks involving object interactions. This…

cs.LG20192 cited

Approximate Inference in Discrete Distributions with Monte Carlo Tree Search and Value Functions

Lars Buesing, Nicolas Heess, Theophane Weber

A plethora of problems in AI, engineering and the sciences are naturally formalized as inference in discrete probabilistic models. Exact inference is often prohibitively expensive,…

cs.LG2019132 cited

Stabilizing Transformers for Reinforcement Learning

Emilio Parisotto, H. Francis Song, Jack W. Rae +10

Owing to their ability to both effectively integrate information over long time horizons and scale to massive amounts of data, self-attention architectures have recently shown brea…

cs.RO20196 cited

Imagined Value Gradients: Model-Based Policy Optimization with Transferable Latent Dynamics Models

Arunkumar Byravan, Jost Tobias Springenberg, Abbas Abdolmaleki +6

Humans are masters at quickly learning many complex tasks, relying on an approximate understanding of the dynamics of their environments. In much the same way, we would like our le…