DeCOM: Decomposed Policy for Constrained Cooperative Multi-Agent Reinforcement Learning
arXiv:2111.05670
Abstract
In recent years, multi-agent reinforcement learning (MARL) has presented impressive performance in various applications. However, physical limitations, budget restrictions, and many other factors usually impose \textit{constraints} on a multi-agent system (MAS), which cannot be handled by traditional MARL frameworks. Specifically, this paper focuses on constrained MASes where agents work \textit{cooperatively} to maximize the expected team-average return under various constraints on expected team-average costs, and develops a \textit{constrained cooperative MARL} framework, named DeCOM, for such MASes. In particular, DeCOM decomposes the policy of each agent into two modules, which empowers information sharing among agents to achieve better cooperation. In addition, with such modularization, the training algorithm of DeCOM separates the original constrained optimization into an unconstrained optimization on reward and a constraints satisfaction problem on costs. DeCOM then iteratively solves these problems in a computationally efficient manner, which makes DeCOM highly scalable. We also provide theoretical guarantees on the convergence of DeCOM's policy update algorithm. Finally, we validate the effectiveness of DeCOM with various types of costs in both toy and large-scale (with 500 agents) environments.
25 pages
References in corpus (19)
- Neural Tangent Kernel: Convergence and Generalization in Neural Networks
- Dota 2 with Large Scale Deep Reinforcement Learning
- Learning a Generic Value-Selection Heuristic Inside a Constraint Programming Solver
- Emergent Tool Use From Multi-Agent Autocurricula
- Learning Attentional Communication for Multi-Agent Cooperation
- Reward Constrained Policy Optimization
- Social Influence as Intrinsic Motivation for Multi-Agent Deep Reinforcement Learning
- Lyapunov-based Safe Policy Optimization for Continuous Control
- Risk-Sensitive and Robust Decision-Making: a CVaR Optimization Approach
- Multi-Robot Collision Avoidance under Uncertainty with Probabilistic Safety Barrier Certificates
- Reinforcement Learning with Convex Constraints
- Learning Safe Multi-Agent Control with Decentralized Neural Barrier Certificates
- Safe Reinforcement Learning via Curriculum Induction
- Variational Policy Gradient Method for Reinforcement Learning with General Utilities
- Value Propagation for Decentralized Networked Deep Multi-agent Reinforcement Learning
- Learning Multi-Agent Coordination for Enhancing Target Coverage in Directional Sensor Networks
- Learning Individually Inferred Communication for Multi-Agent Cooperation
- Constrained Markov Decision Processes via Backward Value Functions
- Sample-Efficient Learning of Stackelberg Equilibria in General-Sum Games