most citedPolicy Gradient Search: Online Planning and Expert Iteration without Search Trees

18 citations · 41 across the 3 of their papers we have counts for

collaborators

5 papers

cs.GT2020

Learning to Resolve Alliance Dilemmas in Many-Player Zero-Sum Games

Edward Hughes, Thomas W. Anthony, Tom Eccles +3

Zero-sum games have long guided artificial intelligence research, since they possess both a rich strategy space of best-responses and a clear evaluation metric. What's more, compet…

cs.GT202016 cited

From Poincaré Recurrence to Convergence in Imperfect Information Games: Finding Equilibrium via Regularization

Julien Perolat, Remi Munos, Jean-Baptiste Lespiau +10

In this paper we investigate the Follow the Regularized Leader dynamics in sequential imperfect information games (IIG). We generalize existing results of Poincaré recurrence from…

cs.LG20207 cited

Smooth markets: A basic mechanism for organizing gradient-based learners

David Balduzzi, Wojciech M Czarnecki, Thomas W Anthony +5

With the success of modern machine learning, it is becoming increasingly important to understand and control how learning algorithms interact. Unfortunately, negative results from…

cs.LG2019

OpenSpiel: A Framework for Reinforcement Learning in Games

Marc Lanctot, Edward Lockhart, Jean-Baptiste Lespiau +24

OpenSpiel is a collection of environments and algorithms for research in general reinforcement learning and search/planning in games. OpenSpiel supports n-player (single- and multi…

cs.LG201918 cited

Policy Gradient Search: Online Planning and Expert Iteration without Search Trees

Thomas Anthony, Robert Nishihara, Philipp Moritz +2

Monte Carlo Tree Search (MCTS) algorithms perform simulation-based search to improve policies online. During search, the simulation policy is adapted to explore the most promising…