activity
20172020
most citedDeepStack: Expert-Level Artificial Intelligence in No-Limit Poker

812 citations · 823 across the 3 of their papers we have counts for

collaborators

7 papers

cs.AI20208 cited

The Advantage Regret-Matching Actor-Critic

Audrūnas Gruslys, Marc Lanctot, Rémi Munos +10

Regret minimization has played a key role in online learning, equilibrium computation in games, and reinforcement learning (RL). In this paper, we describe a general model-free RL…

cs.AI2019

Alternative Function Approximation Parameterizations for Solving Games: An Analysis of -Regression Counterfactual Regret Minimization

Ryan D'Orazio, Dustin Morrill, James R. Wright +1

Function approximation is a powerful approach for structuring large decision problems that has facilitated great achievements in the areas of reinforcement learning and game playin…

cs.LG20193 cited

Bounds for Approximate Regret-Matching Algorithms

Ryan D'Orazio, Dustin Morrill, James R. Wright

A dominant approach to solving large imperfect-information games is Counterfactural Regret Minimization (CFR). In CFR, many regret minimization problems are combined to solve the g…

cs.LG2019

OpenSpiel: A Framework for Reinforcement Learning in Games

Marc Lanctot, Edward Lockhart, Jean-Baptiste Lespiau +24

OpenSpiel is a collection of environments and algorithms for research in general reinforcement learning and search/planning in games. OpenSpiel supports n-player (single- and multi…

cs.LG2019

Neural Replicator Dynamics

Daniel Hennes, Dustin Morrill, Shayegan Omidshafiei +8

Policy gradient and actor-critic algorithms form the basis of many commonly used training techniques in deep reinforcement learning. Using these algorithms in multiagent environmen…

cs.AI2019

Computing Approximate Equilibria in Sequential Adversarial Games by Exploitability Descent

Edward Lockhart, Marc Lanctot, Julien Pérolat +4

In this paper, we present exploitability descent, a new algorithm to compute approximate equilibria in two-player zero-sum extensive-form games with imperfect information, by direc…