Deep Counterfactual Regret Minimization
arXiv:1811.00164
Abstract
Counterfactual Regret Minimization (CFR) is the leading framework for solving large imperfect-information games. It converges to an equilibrium by iteratively traversing the game tree. In order to deal with extremely large games, abstraction is typically applied before running CFR. The abstracted game is solved with tabular CFR, and its solution is mapped back to the full game. This process can be problematic because aspects of abstraction are often manual and domain specific, abstraction algorithms may miss important strategic nuances of the game, and there is a chicken-and-egg problem because determining a good abstraction requires knowledge of the equilibrium of the game. This paper introduces Deep Counterfactual Regret Minimization, a form of CFR that obviates the need for abstraction by instead using deep neural networks to approximate the behavior of CFR in the full game. We show that Deep CFR is principled and achieves strong performance in large poker games. This is the first non-tabular variant of CFR to be successful in large games.
References in corpus (11)
- Adam: A Method for Stochastic Optimization
- Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
- DeepStack: Expert-Level Artificial Intelligence in No-Limit Poker
- Deep Reinforcement Learning from Self-Play in Imperfect-Information Games
- Playing FPS Games with Deep Reinforcement Learning
- A parameter-free hedging algorithm
- Actor-Critic Policy Optimization in Partially Observable Multiagent Environments
- Depth-Limited Solving for Imperfect-Information Games
- Regret Minimization for Partially Observable Deep Reinforcement Learning
- Double Neural Counterfactual Regret Minimization
- Single Deep Counterfactual Regret Minimization
Cited by in corpus (11)
- Mastering the Game of Stratego with Model-Free Multiagent Reinforcement Learning
- RLCard: A Toolkit for Reinforcement Learning in Card Games
- Finding Friend and Foe in Multi-Agent Games
- Neural Replicator Dynamics
- The Advantage Regret-Matching Actor-Critic
- Coordination in Adversarial Sequential Team Games via Multi-Agent Deep Reinforcement Learning
- Multi-Agent Coordination in Adversarial Environments through Signal Mediated Strategies
- Robust Multi-agent Counterfactual Prediction
- Search in Imperfect Information Games
- Deep Synoptic Monte Carlo Planning in Reconnaissance Blind Chess
- Solving zero-sum extensive-form games with arbitrary payoff uncertainty models