Trust Region Policy Optimisation in Multi-Agent Reinforcement Learning
arXiv:2109.11251
Abstract
Trust region methods rigorously enabled reinforcement learning (RL) agents to learn monotonically improving policies, leading to superior performance on a variety of tasks. Unfortunately, when it comes to multi-agent reinforcement learning (MARL), the property of monotonic improvement may not simply apply; this is because agents, even in cooperative games, could have conflicting directions of policy updates. As a result, achieving a guaranteed improvement on the joint policy where each agent acts individually remains an open challenge. In this paper, we extend the theory of trust region learning to MARL. Central to our findings are the multi-agent advantage decomposition lemma and the sequential policy update scheme. Based on these, we develop Heterogeneous-Agent Trust Region Policy Optimisation (HATPRO) and Heterogeneous-Agent Proximal Policy Optimisation (HAPPO) algorithms. Unlike many existing MARL algorithms, HATRPO/HAPPO do not need agents to share parameters, nor do they need any restrictive assumptions on decomposibility of the joint value function. Most importantly, we justify in theory the monotonic improvement property of HATRPO/HAPPO. We evaluate the proposed methods on a series of Multi-Agent MuJoCo and StarCraftII tasks. Results show that HATRPO and HAPPO significantly outperform strong baselines such as IPPO, MAPPO and MADDPG on all tested tasks, therefore establishing a new state of the art.
References in corpus (6)
- Is Independent Learning All You Need in the StarCraft Multi-Agent Challenge?
- Tianshou: a Highly Modularized Deep Reinforcement Learning Library
- Probabilistic Recursive Reasoning for Multi-Agent Reinforcement Learning
- MALib: A Parallel Framework for Population-based Multi-agent Reinforcement Learning
- Settling the Variance of Multi-Agent Policy Gradients
- A Game-Theoretic Approach to Multi-Agent Trust Region Optimization
Cited by in corpus (7)
- FedKL: Tackling Data Heterogeneity in Federated Reinforcement Learning by Penalizing KL Divergence
- Mobility-Aware Decentralized Federated Learning with Joint Optimization of Local Iteration and Leader Selection for Vehicular Networks
- Multi-Agent Constrained Policy Optimisation
- Independent Natural Policy Gradient Always Converges in Markov Potential Games
- Cooperative and Asynchronous Transformer-based Mission Planning for Heterogeneous Teams of Mobile Robots
- Low Variance Trust Region Optimization with Independent Actors and Sequential Updates in Cooperative Multi-agent Reinforcement Learning
- General Automatic Solution Generation of Social Problems