Deep Reinforcement Learning for Swarm Systems
arXiv:1807.06613
Abstract
Recently, deep reinforcement learning (RL) methods have been applied successfully to multi-agent scenarios. Typically, these methods rely on a concatenation of agent states to represent the information content required for decentralized decision making. However, concatenation scales poorly to swarm systems with a large number of homogeneous agents as it does not exploit the fundamental properties inherent to these systems: (i) the agents in the swarm are interchangeable and (ii) the exact number of agents in the swarm is irrelevant. Therefore, we propose a new state representation for deep multi-agent RL based on mean embeddings of distributions. We treat the agents as samples of a distribution and use the empirical mean embedding as input for a decentralized policy. We define different feature spaces of the mean embedding using histograms, radial basis functions and a neural network learned end-to-end. We evaluate the representation on two well known problems from the swarm literature (rendezvous and pursuit evasion), in a globally and locally observable setup. For the local setup we furthermore introduce simple communication protocols. Of all approaches, the mean embedding representation using neural network features enables the richest information exchange between neighboring agents facilitating the development of more complex collective strategies.
31 pages, 12 figures, version 3 (published in JMLR Volume 20)
Cited by in corpus (10)
- Decentralized Multi-Agent Pursuit using Deep Reinforcement Learning
- Autonomous Drone Swarm Navigation and Multi-target Tracking in 3D Environments with Dynamic Obstacles
- Multi-Agent Reinforcement Learning for Dynamic Ocean Monitoring by a Swarm of Buoys
- Online Planning for Multi-UAV Pursuit-Evasion in Unknown Environments Using Deep Reinforcement Learning
- Real-time optimal control of high-dimensional parametrized systems by deep learning-based reduced order models
- Latent feedback control of distributed systems in multiple scenarios through deep learning-based reduced order models
- Context-Aware Deep Q-Network for Decentralized Cooperative Reconnaissance by a Robotic Swarm
- Adaptive Online Distributed Optimal Control of Very-Large-Scale Robotic Systems
- Sim-Env: Decoupling OpenAI Gym Environments from Simulation Models
- Resilient Consensus-based Multi-agent Reinforcement Learning with Function Approximation