"Other-Play" for Zero-Shot Coordination
arXiv:2003.02979
Abstract
We consider the problem of zero-shot coordination - constructing AI agents that can coordinate with novel partners they have not seen before (e.g. humans). Standard Multi-Agent Reinforcement Learning (MARL) methods typically focus on the self-play (SP) setting where agents construct strategies by playing the game with themselves repeatedly. Unfortunately, applying SP naively to the zero-shot coordination problem can produce agents that establish highly specialized conventions that do not carry over to novel partners they have not been trained with. We introduce a novel learning algorithm called other-play (OP), that enhances self-play by looking for more robust strategies, exploiting the presence of known symmetries in the underlying problem. We characterize OP theoretically as well as experimentally. We study the cooperative card game Hanabi and show that OP agents achieve higher scores when paired with independently trained agents. In preliminary results we also show that our OP agents obtains higher average scores when paired with human players, compared to state-of-the-art SP agents.
References in corpus (2)
Cited by in corpus (19)
- Multi-Agent Collaboration via Reward Attribution Decomposition
- Open Problems in Cooperative AI
- Evaluation of Human-AI Teams for Learned and Rule-Based Agents in Hanabi
- Randomized Entity-wise Factorization for Multi-Agent Reinforcement Learning
- Discovering Diverse Multi-Agent Strategic Behavior via Reward Randomization
- Deception in Social Learning: A Multi-Agent Reinforcement Learning Perspective
- Evaluating the Robustness of Collaborative Agents
- Emergent Discrete Communication in Semantic Spaces
- Continuous Coordination As a Realistic Scenario for Lifelong Learning
- Quasi-Equivalence Discovery for Zero-Shot Emergent Communication
- Evaluating the Rainbow DQN Agent in Hanabi with Unseen Partners
- Dynamic population-based meta-learning for multi-agent communication with natural language
- Robust Multi-Agent Reinforcement Learning with Social Empowerment for Coordination and Communication
- Learned Belief Search: Efficiently Improving Policies in Partially Observable Settings
- Exploring Zero-Shot Emergent Communication in Embodied Multi-Agent Populations
- Multi-lingual agents through multi-headed neural networks
- Learning to Communicate with Strangers via Channel Randomisation Methods
- Reinforcement Learning on Human Decision Models for Uniquely Collaborative AI Teammates
- Behaviour-conditioned policies for cooperative reinforcement learning tasks