Actor-Critic Provably Finds Nash Equilibria of Linear-Quadratic Mean-Field Games
arXiv:1910.07498
Abstract
We study discrete-time mean-field Markov games with infinite numbers of agents where each agent aims to minimize its ergodic cost. We consider the setting where the agents have identical linear state transitions and quadratic cost functions, while the aggregated effect of the agents is captured by the population mean of their states, namely, the mean-field state. For such a game, based on the Nash certainty equivalence principle, we provide sufficient conditions for the existence and uniqueness of its Nash equilibrium. Moreover, to find the Nash equilibrium, we propose a mean-field actor-critic algorithm with linear function approximation, which does not require knowing the model of dynamics. Specifically, at each iteration of our algorithm, we use the single-agent actor-critic algorithm to approximately obtain the optimal policy of the each agent given the current mean-field state, and then update the mean-field state. In particular, we prove that our algorithm converges to the Nash equilibrium at a linear rate. To the best of our knowledge, this is the first success of applying model-free reinforcement learning with function approximation to discrete-time mean-field Markov games with provable non-asymptotic global convergence guarantees.
References in corpus (7)
- DeepStack: Expert-Level Artificial Intelligence in No-Limit Poker
- Safe, Multi-Agent, Reinforcement Learning for Autonomous Driving
- Finite-Sample Analysis of Proximal Gradient TD Algorithms
- Value Function Approximation in Zero-Sum Markov Games
- Least-Squares Temporal Difference Learning for the Linear Quadratic Regulator
- On the Global Convergence of Actor-Critic: A Case for Linear Quadratic Regulator with Ergodic Cost
- Linear quadratic mean field games: Asymptotic solvability and relation to the fixed point approach
Cited by in corpus (8)
- Game-Theoretic Multiagent Reinforcement Learning
- Single-Timescale Actor-Critic Provably Finds Globally Optimal Policy
- Entropy Regularization for Mean Field Games with Learning
- Non-asymptotic Convergence of Adam-type Reinforcement Learning Algorithms under Markovian Sampling
- Provable Fictitious Play for General Mean-Field Games
- Alternating the Population and Control Neural Networks to Solve High-Dimensional Stochastic Mean-Field Games
- Global Convergence of Policy Gradient for Linear-Quadratic Mean-Field Control/Game in Continuous Time
- Model Free Reinforcement Learning Algorithm for Stationary Mean field Equilibrium for Multiple Types of Agents