activity
20172022
most citeddm_control: Software and Tasks for Continuous Control

198 citations · 282 across the 8 of their papers we have counts for

collaborators

11 papers

cs.MA2022

Developing, Evaluating and Scaling Learning Agents in Multi-Agent Environments

Ian Gemp, Thomas Anthony, Yoram Bachrach +24

The Game Theory & Multi-Agent team at DeepMind studies several aspects of multi-agent learning ranging from computing approximations to fundamental concepts in game theory to simul…

cs.LG2022

Revisiting Gaussian mixture critics in off-policy reinforcement learning: a sample-based approach

Bobak Shahriari, Abbas Abdolmaleki, Arunkumar Byravan +6

Actor-critic algorithms that make use of distributional policy evaluation have frequently been shown to outperform their non-distributional counterparts on many challenging control…

cs.AI20221 cited

NeuPL: Neural Population Learning

Siqi Liu, Luke Marris, Daniel Hennes +3

Learning in strategy games (e.g. StarCraft, poker) requires the discovery of diverse policies. This is often achieved by iteratively training new policies against existing ones, gr…

cs.AI2021

Pick Your Battles: Interaction Graphs as Population-Level Objectives for Strategic Diversity

Marta Garnelo, Wojciech Marian Czarnecki, Siqi Liu +5

Strategic diversity is often essential in games: in multi-player games, for example, evaluating a player against a diverse set of strategies will yield a more accurate estimate of…

cs.DC20215 cited

Launchpad: A Programming Model for Distributed Machine Learning Research

Fan Yang, Gabriel Barth-Maron, Piotr Stańczyk +5

A major driver behind the success of modern machine learning algorithms has been their ability to process ever-larger amounts of data. As a result, the use of distributed systems i…

cs.AI202112 cited

From Motor Control to Team Play in Simulated Humanoid Football

Siqi Liu, Guy Lever, Zhe Wang +19

Intelligent behaviour in the physical world exhibits structure at multiple spatial and temporal scales. Although movements are ultimately executed at the level of instantaneous mus…