Publications (15)
Alternative Function Approximation Parameterizations for Solving Games: An Analysis of -Regression Counterfactual Regret Minimization
Ryan D'Orazio, Dustin Morrill, James R. Wright +1
Function approximation is a powerful approach for structuring large decision problems that has facilitated great achievements in the areas of reinforcement learning and game playin…
Solving Games with Functional Regret Estimation
Kevin Waugh, Dustin Morrill, J. Andrew Bagnell +1
We propose a novel online learning method for minimizing regret in large extensive-form games. The approach learns a function approximator online to estimate the regret for choosin…
Interpolating Between Softmax Policy Gradient and Neural Replicator Dynamics with Capped Implicit Exploration
Dustin Morrill, Esra'a Saleh, Michael Bowling +1
Neural replicator dynamics (NeuRD) is an alternative to the foundational softmax policy gradient (SPG) algorithm motivated by online learning and evolutionary game theory. The NeuR…
Neural Replicator Dynamics
Daniel Hennes, Dustin Morrill, Shayegan Omidshafiei +8
Policy gradient and actor-critic algorithms form the basis of many commonly used training techniques in deep reinforcement learning. Using these algorithms in multiagent environmen…
Hindsight and Sequential Rationality of Correlated Play
Dustin Morrill, Ryan D'Orazio, Reca Sarfati +4
Driven by recent successes in two-player, zero-sum game solving and playing, artificial intelligence work on games has increasingly focused on algorithms that produce equilibrium-b…
Efficient Deviation Types and Learning for Hindsight Rationality in Extensive-Form Games: Corrections
Dustin Morrill, Ryan D'Orazio, Marc Lanctot +3
Hindsight rationality is an approach to playing general-sum games that prescribes no-regret learning dynamics for individual agents with respect to a set of deviations, and further…
DeepStack: Expert-Level Artificial Intelligence in No-Limit Poker
Matej MoravÄÃk, Martin Schmid, Neil Burch +7
Artificial intelligence has seen several breakthroughs in recent years, with games often serving as milestones. A common feature of these games is that players have perfect informa…
OpenSpiel: A Framework for Reinforcement Learning in Games
Marc Lanctot, Edward Lockhart, Jean-Baptiste Lespiau +24
OpenSpiel is a collection of environments and algorithms for research in general reinforcement learning and search/planning in games. OpenSpiel supports n-player (single- and multi…
Learning to Be Cautious
Montaser Mohammedalamen, Dustin Morrill, Alexander Sieusahai +2
A key challenge in the field of reinforcement learning is to develop agents that behave cautiously in novel situations. It is generally impossible to anticipate all situations that…
Composing Efficient, Robust Tests for Policy Selection
Dustin Morrill, Thomas J. Walsh, Daniel Hernandez +2
Modern reinforcement learning systems produce many high-quality policies throughout the learning process. However, to choose which policy to actually deploy in the real world, they…
The Advantage Regret-Matching Actor-Critic
Audrūnas Gruslys, Marc Lanctot, Rémi Munos +10
Regret minimization has played a key role in online learning, equilibrium computation in games, and reinforcement learning (RL). In this paper, we describe a general model-free RL…
Bounds for Approximate Regret-Matching Algorithms
Ryan D'Orazio, Dustin Morrill, James R. Wright
A dominant approach to solving large imperfect-information games is Counterfactural Regret Minimization (CFR). In CFR, many regret minimization problems are combined to solve the g…
Computing Approximate Equilibria in Sequential Adversarial Games by Exploitability Descent
Edward Lockhart, Marc Lanctot, Julien Pérolat +4
In this paper, we present exploitability descent, a new algorithm to compute approximate equilibria in two-player zero-sum extensive-form games with imperfect information, by direc…
The Partially Observable History Process
Dustin Morrill, Amy R. Greenwald, Michael Bowling
We introduce the partially observable history process (POHP) formalism for reinforcement learning. POHP centers around the actions and observations of a single agent and abstracts…
Efficient Deviation Types and Learning for Hindsight Rationality in Extensive-Form Games
Dustin Morrill, Ryan D'Orazio, Marc Lanctot +3
Hindsight rationality is an approach to playing general-sum games that prescribes no-regret learning dynamics for individual agents with respect to a set of deviations, and further…