-Learning: A Collaborative Distributed Strategy for Multi-Agent Reinforcement Learning Through Consensus + Innovations
arXiv:1205.0047 · doi:10.1109/TSP.2013.2241057
Abstract
The paper considers a class of multi-agent Markov decision processes (MDPs), in which the network agents respond differently (as manifested by the instantaneous one-stage random costs) to a global controlled state and the control actions of a remote controller. The paper investigates a distributed reinforcement learning setup with no prior information on the global state transition and local agent cost statistics. Specifically, with the agents' objective consisting of minimizing a network-averaged infinite horizon discounted cost, the paper proposes a distributed version of -learning, -learning, in which the network agents collaborate by means of local processing and mutual information exchange over a sparse (possibly stochastic) communication network to achieve the network goal. Under the assumption that each agent is only aware of its local online cost data and the inter-agent communication network is \emph{weakly} connected, the proposed distributed scheme is almost surely (a.s.) shown to yield asymptotically the desired value function and the optimal stationary control policy at each network agent. The analytical techniques developed in the paper to address the mixed time-scale stochastic dynamics of the \emph{consensus + innovations} form, which arise as a result of the proposed interactive distributed scheme, are of independent interest.
Submitted to the IEEE Transactions on Signal Processing, 33 pages
References in corpus (4)
- Gossip Algorithms for Distributed Signal Processing
- The Communicative Multiagent Team Decision Problem: Analyzing Teamwork Theories and Models
- Convergence Rate Analysis of Distributed Gossip (Linear Parameter) Estimation: Fundamental Limits and Tradeoffs
- Cooperative Convex Optimization in Networked Systems: Augmented Lagrangian Algorithms with Directed Gossip Communication
Cited by in corpus (36)
- Towards Massive Machine Type Communications in Ultra-Dense Cellular IoT Networks: Current Issues and Machine Learning-Assisted Solutions
- Fully Decentralized Multi-Agent Reinforcement Learning with Networked Agents
- Optimization for Reinforcement Learning: From Single Agent to Cooperative Agents
- Towards Understanding Asynchronous Advantage Actor-critic: Convergence and Linear Speedup
- Diff-DAC: Distributed Actor-Critic for Average Multitask Deep Reinforcement Learning
- Scalable Multi-Agent Reinforcement Learning for Networked Systems with Average Reward
- Minimax Robust Detection: Classic Results and Recent Advances
- Byzantine-Resilient Decentralized TD Learning with Linear Function Approximation
- Decentralized Multi-Agent Reinforcement Learning with Networked Agents: Recent Advances
- A Discrete-Time Switching System Analysis of Q-learning
- Applications of Multi-Agent Reinforcement Learning in Future Internet: A Comprehensive Survey
- F2A2: Flexible Fully-decentralized Approximate Actor-critic for Cooperative Multi-agent Reinforcement Learning
- Multi-Agent Trust Region Policy Optimization
- Multi-Timescale Ensemble Q-learning for Markov Decision Process Policy Optimization
- Asynchronous Distributed Reinforcement Learning for LQR Control via Zeroth-Order Block Coordinate Descent
- A Multi-Agent Off-Policy Actor-Critic Algorithm for Distributed Reinforcement Learning
- Primal-Dual Distributed Temporal Difference Learning
- Clustering with Distributed Data
- A Decentralized Policy Gradient Approach to Multi-task Reinforcement Learning
- Coach-Player Multi-Agent Reinforcement Learning for Dynamic Team Composition
- MARL with General Utilities via Decentralized Shadow Reward Actor-Critic
- Provably Efficient Cooperative Multi-Agent Reinforcement Learning with Function Approximation
- Regret Bounds for Decentralized Learning in Cooperative Multi-Agent Dynamical Systems
- Finite-Sample Analysis For Decentralized Batch Multi-Agent Reinforcement Learning With Networked Agents
- Communication-Efficient Policy Gradient Methods for Distributed Reinforcement Learning
- A Reinforcement Learning Framework for Sequencing Multi-Robot Behaviors
- Learning to Gather without Communication
- A Communication-Efficient Multi-Agent Actor-Critic Algorithm for Distributed Reinforcement Learning
- Towards Resilience for Multi-Agent -Learning
- Fully Distributed Actor-Critic Architecture for Multitask Deep Reinforcement Learning
- Programming and Deployment of Autonomous Swarms using Multi-Agent Reinforcement Learning
- Exploiting Fast Decaying and Locality in Multi-Agent MDP with Tree Dependence Structure
- Distributed Safe Learning using an Invariance-based Safety Framework
- A Law of Iterated Logarithm for Multi-Agent Reinforcement Learning
- Resilient Consensus-based Multi-agent Reinforcement Learning with Function Approximation
- Distributed Policy Evaluation Under Multiple Behavior Strategies