activity
20132017
most citedMastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

1.1k citations · 1.5k across the 4 of their papers we have counts for

collaborators

8 papers

cs.AI20171.1k cited

Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

David Silver, Thomas Hubert, Julian Schrittwieser +10

The game of chess is the most widely-studied domain in the history of artificial intelligence. The strongest programs are based on a combination of sophisticated search techniques,…

cs.AI2017142 cited

A Unified Game-Theoretic Approach to Multiagent Reinforcement Learning

Marc Lanctot, Vinicius Zambaldi, Audrunas Gruslys +5

To achieve general intelligence, agents must learn how to interact with others in a shared environment: this is the challenge of multiagent reinforcement learning (MARL). The simpl…

cs.GT2017

Symmetric Decomposition of Asymmetric Games

Karl Tuyls, Julien Perolat, Marc Lanctot +6

We introduce new theoretical insights into two-population asymmetric games allowing for an elegant symmetric decomposition into two single population symmetric games. Specifically,…

cs.AI2017

Value-Decomposition Networks For Cooperative Multi-Agent Learning

Peter Sunehag, Guy Lever, Audrunas Gruslys +8

We study the problem of cooperative multi-agent reinforcement learning with a single joint reward signal. This class of learning problems is difficult because of the often large co…

cs.MA2017277 cited

Multi-agent Reinforcement Learning in Sequential Social Dilemmas

Joel Z. Leibo, Vinicius Zambaldi, Marc Lanctot +2

Matrix games like Prisoner's Dilemma have guided research on social dilemmas for decades. However, they necessarily treat the choice to cooperate or defect as an atomic action. In…

cs.NE2016

Memory-Efficient Backpropagation Through Time

Audrūnas Gruslys, Remi Munos, Ivo Danihelka +2

We propose a novel approach to reduce memory consumption of the backpropagation through time (BPTT) algorithm when training recurrent neural networks (RNNs). Our approach uses dyna…