On the interaction between supervision and self-play in emergent communication
arXiv:2002.01093
Abstract
A promising approach for teaching artificial agents to use natural language involves using human-in-the-loop training. However, recent work suggests that current machine learning methods are too data inefficient to be trained in this way from scratch. In this paper, we investigate the relationship between two categories of learning signals with the ultimate goal of improving sample efficiency: imitating human language data via supervised learning, and maximizing reward in a simulated multi-agent environment via self-play (as done in emergent communication), and introduce the term supervised self-play (S2P) for algorithms using both of these signals. We find that first training agents via supervised learning on human data followed by self-play outperforms the converse, suggesting that it is not beneficial to emerge languages from scratch. We then empirically investigate various S2P schedules that begin with supervised learning in two environments: a Lewis signaling game with symbolic inputs, and an image-based referential game with natural language descriptions. Lastly, we introduce population based approaches to S2P, which further improves the performance over single-agent methods.
The first two authors contributed equally. Accepted at ICLR 2020
References in corpus (2)
Cited by in corpus (15)
- Emergent Multi-Agent Communication in the Deep Learning Era
- Emergent Language: A Survey and Taxonomy
- Emergent Discrete Communication in Semantic Spaces
- Quasi-Equivalence Discovery for Zero-Shot Emergent Communication
- Dynamic population-based meta-learning for multi-agent communication with natural language
- Bridging the Imitation Gap by Adaptive Insubordination
- Iterated learning for emergent systematicity in VQA
- Exploring Zero-Shot Emergent Communication in Embodied Multi-Agent Populations
- Structural Inductive Biases in Emergent Communication
- Learning to Communicate with Strangers via Channel Randomisation Methods
- SocialAI: Benchmarking Socio-Cognitive Abilities in Deep Reinforcement Learning Agents
- Self-play for Data Efficient Language Acquisition
- The Emergence of the Shape Bias Results from Communicative Efficiency
- Targeted Data Acquisition for Evolving Negotiation Agents
- Supervised Seeded Iterated Learning for Interactive Language Learning