132 citations · 207 across the 4 of their papers we have counts for
9 papers
From Motor Control to Team Play in Simulated Humanoid Football
Siqi Liu, Guy Lever, Zhe Wang +19
Intelligent behaviour in the physical world exhibits structure at multiple spatial and temporal scales. Although movements are ultimately executed at the level of instantaneous mus…
A Distributional View on Multi-Objective Policy Optimization
Abbas Abdolmaleki, Sandy H. Huang, Leonard Hasenclever +7
Many real-world problems require trading off multiple competing objectives. However, these objectives are often in different units and/or scales, which can make it challenging for…
Stabilizing Transformers for Reinforcement Learning
Emilio Parisotto, H. Francis Song, Jack W. Rae +10
Owing to their ability to both effectively integrate information over long time horizons and scale to massive amounts of data, self-attention architectures have recently shown brea…
V-MPO: On-Policy Maximum a Posteriori Policy Optimization for Discrete and Continuous Control
H. Francis Song, Abbas Abdolmaleki, Jost Tobias Springenberg +11
Some of the most successful applications of deep reinforcement learning to challenging domains in discrete and continuous control have used policy gradient methods in the on-policy…
The Hanabi Challenge: A New Frontier for AI Research
Nolan Bard, Jakob N. Foerster, Sarath Chandar +12
From the early days of computing, games have been important testbeds for studying how well machines can do sophisticated decision making. In recent years, machine learning has made…
Bayesian Action Decoder for Deep Multi-Agent Reinforcement Learning
Jakob N. Foerster, Francis Song, Edward Hughes +5
When observing the actions of others, humans make inferences about why they acted as they did, and what this implies about the world; humans also use the fact that their actions wi…