Double Neural Counterfactual Regret Minimization
arXiv:1812.10607
Abstract
Counterfactual Regret Minimization (CRF) is a fundamental and effective technique for solving Imperfect Information Games (IIG). However, the original CRF algorithm only works for discrete state and action spaces, and the resulting strategy is maintained as a tabular representation. Such tabular representation limits the method from being directly applied to large games and continuing to improve from a poor strategy profile. In this paper, we propose a double neural representation for the imperfect information games, where one neural network represents the cumulative regret, and the other represents the average strategy. Furthermore, we adopt the counterfactual regret minimization algorithm to optimize this double neural representation. To make neural learning efficient, we also developed several novel techniques including a robust sampling method, mini-batch Monte Carlo Counterfactual Regret Minimization (MCCFR) and Monte Carlo Counterfactual Regret Minimization Plus (MCCFR+) which may be of independent interests. Experimentally, we demonstrate that the proposed double neural algorithm converges significantly better than the reinforcement learning counterpart.
References in corpus (4)
Cited by in corpus (10)
- Deep Counterfactual Regret Minimization
- TLeague: A Framework for Competitive Self-Play based Distributed Multi-Agent Reinforcement Learning
- Algorithms in Multi-Agent Systems: A Holistic Perspective from Reinforcement Learning and Game Theory
- The Advantage Regret-Matching Actor-Critic
- Solving imperfect-information games via exponential counterfactual regret minimization
- Model-free Neural Counterfactual Regret Minimization with Bootstrap Learning
- Equivalence Analysis between Counterfactual Regret Minimization and Online Mirror Descent
- Search in Imperfect Information Games
- Solving zero-sum extensive-form games with arbitrary payoff uncertainty models
- RLCFR: Minimize Counterfactual Regret by Deep Reinforcement Learning