Sample and Communication-Efficient Decentralized Actor-Critic Algorithms with Finite-Time Analysis
arXiv:2109.03699
Abstract
Actor-critic (AC) algorithms have been widely adopted in decentralized multi-agent systems to learn the optimal joint control policy. However, existing decentralized AC algorithms either do not preserve the privacy of agents or are not sample and communication-efficient. In this work, we develop two decentralized AC and natural AC (NAC) algorithms that are private, and sample and communication-efficient. In both algorithms, agents share noisy information to preserve privacy and adopt mini-batch updates to improve sample and communication efficiency. Particularly for decentralized NAC, we develop a decentralized Markovian SGD algorithm with an adaptive mini-batch size to efficiently compute the natural policy gradient. Under Markovian sampling and linear function approximation, we prove the proposed decentralized AC and NAC algorithms achieve the state-of-the-art sample complexities and , respectively, and the same small communication complexity . Numerical experiments demonstrate that the proposed algorithms achieve lower sample and communication complexities than the existing decentralized AC algorithm.
40 pages, 2 figures
References in corpus (9)
- Intelligent Electric Vehicle Charging Recommendation Based on Multi-Agent Reinforcement Learning
- Finite-Time Analysis of Distributed TD(0) with Linear Function Approximation for Multi-Agent Reinforcement Learning
- Two Time-scale Off-Policy TD Learning: Non-asymptotic Analysis over Markovian Samples
- Non-asymptotic Convergence Analysis of Two Time-scale (Natural) Actor-Critic Algorithms
- An Improved Analysis of (Variance-Reduced) Policy Gradient and Natural Policy Gradient Methods
- Distributed TD(0) with Almost No Communication
- Multi-Agent Off-Policy TD Learning: Finite-Time Analysis with Near-Optimal Sample Complexity and Communication Complexity
- Shaping Advice in Deep Multi-Agent Reinforcement Learning
- Model Free Reinforcement Learning Algorithm for Stationary Mean field Equilibrium for Multiple Types of Agents