Model-Free Mean-Field Reinforcement Learning: Mean-Field MDP and Mean-Field Q-Learning
arXiv:1910.12802
Abstract
We study infinite horizon discounted Mean Field Control (MFC) problems with common noise through the lens of Mean Field Markov Decision Processes (MFMDP). We allow the agents to use actions that are randomized not only at the individual level but also at the level of the population. This common randomization allows us to establish connections between both closed-loop and open-loop policies for MFC and Markov policies for the MFMDP. In particular, we show that there exists an optimal closed-loop policy for the original MFC. Building on this framework and the notion of state-action value function, we then propose reinforcement learning (RL) methods for such problems, by adapting existing tabular and deep RL methods to the mean-field setting. The main difficulty is the treatment of the population state, which is an input of the policy and the value function. We provide convergence guarantees for tabular algorithms based on discretizations of the simplex. Neural network based algorithms are more suitable for continuous spaces and allow us to avoid discretizing the mean field state space. Numerical examples are provided.
References in corpus (7)
- A Machine Learning Framework for Solving High-Dimensional Mean Field Game and Mean Field Control Problems
- Fictitious Play for Mean Field Games: Continuous Time Analysis and Applications
- Linear-Quadratic Mean-Field Reinforcement Learning: Convergence of Policy Gradient Methods
- Actor-Critic Provably Finds Nash Equilibria of Linear-Quadratic Mean-Field Games
- Dynamic Programming Principles for Mean-Field Controls with Learning
- Reinforcement Learning for Mean Field Game
- Mean-Field Controls with Q-learning for Cooperative MARL: Convergence and Complexity Analysis
Cited by in corpus (11)
- Game-Theoretic Multiagent Reinforcement Learning
- Neural networks-based algorithms for stochastic control and PDEs in finance
- Dynamic Programming Principles for Mean-Field Controls with Learning
- Mean Field Markov Decision Processes
- Learning in Discounted-cost and Average-cost Mean-field Games
- On the Approximation of Cooperative Heterogeneous Multi-Agent Reinforcement Learning (MARL) using Mean Field Control (MFC)
- Exploration noise for learning linear-quadratic mean field games
- Permutation Invariant Policy Optimization for Mean-Field Multi-Agent Reinforcement Learning: A Principled Approach
- Adaptive Online Distributed Optimal Control of Very-Large-Scale Robotic Systems
- Model Free Reinforcement Learning Algorithm for Stationary Mean field Equilibrium for Multiple Types of Agents
- Policy Optimization for Linear-Quadratic Zero-Sum Mean-Field Type Games