PAC Guarantees for Cooperative Multi-Agent Reinforcement Learning with Restricted Communication
arXiv:1905.09951
Abstract
We develop model free PAC performance guarantees for multiple concurrent MDPs, extending recent works where a single learner interacts with multiple non-interacting agents in a noise free environment. Our framework allows noisy and resource limited communication between agents, and develops novel PAC guarantees in this extended setting. By allowing communication between the agents themselves, we suggest improved PAC-exploration algorithms that can overcome the communication noise and lead to improved sample complexity bounds. We provide a theoretically motivated algorithm that optimally combines information from the resource limited agents, thereby analyzing the interaction between noise and communication constraints that are ubiquitous in real-world systems. We present empirical results for a simple task that supports our theoretical formulations and improve upon naive information fusion methods.
References in corpus (7)
- Learning Attentional Communication for Multi-Agent Cooperation
- Deep Decentralized Multi-task Multi-Agent Reinforcement Learning under Partial Observability
- Emergent Complexity via Multi-Agent Competition
- Unifying PAC and Regret: Uniform PAC Bounds for Episodic Reinforcement Learning
- Decentralized Cooperative Stochastic Bandits
- Scalable Coordinated Exploration in Concurrent Reinforcement Learning
- Concurrent Meta Reinforcement Learning