Neural Replicator Dynamics
arXiv:1906.00190
Abstract
Policy gradient and actor-critic algorithms form the basis of many commonly used training techniques in deep reinforcement learning. Using these algorithms in multiagent environments poses problems such as nonstationarity and instability. In this paper, we first demonstrate that standard softmax-based policy gradient can be prone to poor performance in the presence of even the most benign nonstationarity. By contrast, it is known that the replicator dynamics, a well-studied model from evolutionary game theory, eliminates dominated strategies and exhibits convergence of the time-averaged trajectories to interior Nash equilibria in zero-sum games. Thus, using the replicator dynamics as a foundation, we derive an elegant one-line change to policy gradient methods that simply bypasses the gradient step through the softmax, yielding a new algorithm titled Neural Replicator Dynamics (NeuRD). NeuRD reduces to the exponential weights/Hedge algorithm in the single-state all-actions case. Additionally, NeuRD has formal equivalence to softmax counterfactual regret minimization, which guarantees convergence in the sequential tabular case. Importantly, our algorithm provides a straightforward way of extending the replicator dynamics to the function approximation setting. Empirical results show that NeuRD quickly adapts to nonstationarities, outperforming policy gradient significantly in both tabular and function approximation settings, when evaluated on the standard imperfect information benchmarks of Kuhn Poker, Leduc Poker, and Goofspiel.
References in corpus (18)
- Continuous control with deep reinforcement learning
- Trust Region Policy Optimization
- Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments
- DeepStack: Expert-Level Artificial Intelligence in No-Limit Poker
- IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures
- Counterfactual Multi-Agent Policy Gradients
- Optimizing Neural Networks with Kronecker-factored Approximate Curvature
- The Hanabi Challenge: A New Frontier for AI Research
- Continuous Adaptation via Meta-Learning in Nonstationary and Competitive Environments
- Bayes' Bluff: Opponent Modelling in Poker
- Emergent Complexity via Multi-Agent Competition
- Deep Reinforcement Learning from Self-Play in Imperfect-Information Games
- OpenSpiel: A Framework for Reinforcement Learning in Games
- Actor-Critic Policy Optimization in Partially Observable Multiagent Environments
- Deep Counterfactual Regret Minimization
- From Poincaré Recurrence to Convergence in Imperfect Information Games: Finding Equilibrium via Regularization
- Bounds for Approximate Regret-Matching Algorithms
- Alternative Function Approximation Parameterizations for Solving Games: An Analysis of -Regression Counterfactual Regret Minimization
Cited by in corpus (8)
- A Survey and Critique of Multiagent Deep Reinforcement Learning
- Game-Theoretic Multiagent Reinforcement Learning
- Reinforcement learning for pursuit and evasion of microswimmers at low Reynolds number
- SEED RL: Scalable and Efficient Deep-RL with Accelerated Central Inference
- DREAM: Deep Regret minimization with Advantage baselines and Model-free learning
- Independent Natural Policy Gradient Always Converges in Markov Potential Games
- Bounds for Approximate Regret-Matching Algorithms
- Alternative Function Approximation Parameterizations for Solving Games: An Analysis of -Regression Counterfactual Regret Minimization