MAVEN: Multi-Agent Variational Exploration
arXiv:1910.07483
Abstract
Centralised training with decentralised execution is an important setting for cooperative deep multi-agent reinforcement learning due to communication constraints during execution and computational tractability in training. In this paper, we analyse value-based methods that are known to have superior performance in complex environments [43]. We specifically focus on QMIX [40], the current state-of-the-art in this domain. We show that the representational constraints on the joint action-values introduced by QMIX and similar methods lead to provably poor exploration and suboptimality. Furthermore, we propose a novel approach called MAVEN that hybridises value and policy-based methods by introducing a latent space for hierarchical control. The value-based agents condition their behaviour on the shared latent variable controlled by a hierarchical policy. This allows MAVEN to achieve committed, temporally extended exploration, which is key to solving complex multi-agent tasks. Our experimental results show that MAVEN achieves significant performance improvements on the challenging SMAC domain [43].
Cited by in corpus (19)
- Is Independent Learning All You Need in the StarCraft Multi-Agent Challenge?
- SMARTS: Scalable Multi-Agent Reinforcement Learning Training School for Autonomous Driving
- RODE: Learning Roles to Decompose Multi-Agent Tasks
- UPDeT: Universal Multi-agent Reinforcement Learning via Policy Decoupling with Transformers
- Q-value Path Decomposition for Deep Multiagent Reinforcement Learning
- QTRAN++: Improved Value Transformation for Cooperative Multi-Agent Reinforcement Learning
- RMIX: Learning Risk-Sensitive Policies for Cooperative Reinforcement Learning Agents
- Discovering Diverse Multi-Agent Strategic Behavior via Reward Randomization
- Multi-Agent Determinantal Q-Learning
- A Maximum Mutual Information Framework for Multi-Agent Reinforcement Learning
- Credit Assignment with Meta-Policy Gradient for Multi-Agent Reinforcement Learning
- Model based Multi-agent Reinforcement Learning with Tensor Decompositions
- Skill Discovery of Coordination in Multi-agent Reinforcement Learning
- Adaptable Agent Populations via a Generative Model of Policies
- QVMix and QVMix-Max: Extending the Deep Quality-Value Family of Algorithms to Cooperative Multi-Agent Reinforcement Learning
- Reinforcement Learning in Factored Action Spaces using Tensor Decompositions
- Coordinated Proximal Policy Optimization
- Regularize! Don't Mix: Multi-Agent Reinforcement Learning without Explicit Centralized Structures
- DSDF: An approach to handle stochastic agents in collaborative multi-agent reinforcement learning