Deep Reinforcement Learning from Self-Play in Imperfect-Information Games
arXiv:1603.01121
Abstract
Many real-world applications can be described as large-scale games of imperfect information. To deal with these challenging domains, prior work has focused on computing Nash equilibria in a handcrafted abstraction of the domain. In this paper we introduce the first scalable end-to-end approach to learning approximate Nash equilibria without prior domain knowledge. Our method combines fictitious self-play with deep reinforcement learning. When applied to Leduc poker, Neural Fictitious Self-Play (NFSP) approached a Nash equilibrium, whereas common reinforcement learning methods diverged. In Limit Texas Holdem, a poker game of real-world scale, NFSP learnt a strategy that approached the performance of state-of-the-art, superhuman algorithms based on significant domain expertise.
updated version, incorporating conference feedback
References in corpus (3)
Cited by in corpus (61)
- Dota 2 with Large Scale Deep Reinforcement Learning
- DeepStack: Expert-Level Artificial Intelligence in No-Limit Poker
- Human-level performance in first-person multiplayer games with population-based deep reinforcement learning
- A Survey and Critique of Multiagent Deep Reinforcement Learning
- Game-Theoretic Multiagent Reinforcement Learning
- Emergent Complexity via Multi-Agent Competition
- Maintaining cooperation in complex social dilemmas using deep reinforcement learning
- Finding Effective Security Strategies through Reinforcement Learning and Self-Play
- Open-Ended Learning Leads to Generally Capable Agents
- RLCard: A Toolkit for Reinforcement Learning in Card Games
- Deep Counterfactual Regret Minimization
- Consequentialist conditional cooperation in social dilemmas with imperfect information
- Actor-Critic Provably Finds Nash Equilibria of Linear-Quadratic Mean-Field Games
- Learning to Play No-Press Diplomacy with Best Response Policy Iteration
- Applications of Deep Reinforcement Learning in Communications and Networking: A Survey
- From Poincaré Recurrence to Convergence in Imperfect Information Games: Finding Equilibrium via Regularization
- Double Neural Counterfactual Regret Minimization
- Single Deep Counterfactual Regret Minimization
- Algorithms in Multi-Agent Systems: A Holistic Perspective from Reinforcement Learning and Game Theory
- Neural Replicator Dynamics
- DREAM: Deep Regret minimization with Advantage baselines and Model-free learning
- Pipeline PSRO: A Scalable Approach for Finding Approximate Nash Equilibria in Large Games
- Deep Fictitious Play for Stochastic Differential Games
- Human-Level Performance in No-Press Diplomacy via Equilibrium Search
- Arena: A General Evaluation Platform and Building Toolkit for Multi-Agent Intelligence
- Robust Opponent Modeling via Adversarial Ensemble Reinforcement Learning in Asymmetric Imperfect-Information Games
- Fever Basketball: A Complex, Flexible, and Asynchronized Sports Game Environment for Multi-agent Reinforcement Learning
- Joint Policy Search for Multi-agent Collaboration with Imperfect Information
- Global Convergence of Policy Gradient for Linear-Quadratic Mean-Field Control/Game in Continuous Time
- Deep Reinforcement Learning in Quantitative Algorithmic Trading: A Review
- Application of Self-Play Reinforcement Learning to a Four-Player Game of Imperfect Information
- Neural Fictitious Self-Play on ELF Mini-RTS
- Convergence Analysis of Gradient-Based Learning with Non-Uniform Learning Rates in Non-Cooperative Multi-Agent Settings
- Empirical Analysis of Fictitious Play for Nash Equilibrium Computation in Multiplayer Games
- XDO: A Double Oracle Algorithm for Extensive-Form Games
- Unsupervised Visual Attention and Invariance for Reinforcement Learning
- Independent Natural Policy Gradient Always Converges in Markov Potential Games
- Optimize Neural Fictitious Self-Play in Regret Minimization Thinking
- Gamifying the Vehicle Routing Problem with Stochastic Requests
- Non-cooperative Multi-agent Systems with Exploring Agents
- Last-iterate Convergence in Extensive-Form Games
- Coordination in Adversarial Sequential Team Games via Multi-Agent Deep Reinforcement Learning
- DRL: Deep Reinforcement Learning for Intelligent Robot Control -- Concept, Literature, and Future
- Deep Latent Competition: Learning to Race Using Visual Control Policies in Latent Space
- Multi-Agent Coordination in Adversarial Environments through Signal Mediated Strategies
- Solving Large-Scale Extensive-Form Network Security Games via Neural Fictitious Self-Play
- Approximate exploitability: Learning a best response in large games
- Robust Multi-agent Counterfactual Prediction
- Optimistic Distributionally Robust Policy Optimization
- Alternative Function Approximation Parameterizations for Solving Games: An Analysis of -Regression Counterfactual Regret Minimization
- Latent Dirichlet Allocation for Internet Price War
- Improving Fictitious Play Reinforcement Learning with Expanding Models
- Interactive Agent Modeling by Learning to Probe
- Active Perception in Adversarial Scenarios using Maximum Entropy Deep Reinforcement Learning
- Efficient Competitive Self-Play Policy Optimization
- Improved Robustness and Safety for Autonomous Vehicle Control with Adversarial Reinforcement Learning
- Learning in the Machine: the Symmetries of the Deep Learning Channel
- Multi-agent Reinforcement Learning in OpenSpiel: A Reproduction Report
- On Solving Cooperative MARL Problems with a Few Good Experiences
- A Dynamics Perspective of Pursuit-Evasion Games of Intelligent Agents with the Ability to Learn
- Value Variance Minimization for Learning Approximate Equilibrium in Aggregation Systems