Finite-Time Error Bounds for Distributed Linear Stochastic Approximation
arXiv:2111.12665
Abstract
This paper considers a novel multi-agent linear stochastic approximation algorithm driven by Markovian noise and general consensus-type interaction, in which each agent evolves according to its local stochastic approximation process which depends on the information from its neighbors. The interconnection structure among the agents is described by a time-varying directed graph. While the convergence of consensus-based stochastic approximation algorithms when the interconnection among the agents is described by doubly stochastic matrices (at least in expectation) has been studied, less is known about the case when the interconnection matrix is simply stochastic. For any uniformly strongly connected graph sequences whose associated interaction matrices are stochastic, the paper derives finite-time bounds on the mean-square error, defined as the deviation of the output of the algorithm from the unique equilibrium point of the associated ordinary differential equation. For the case of interconnection matrices being stochastic, the equilibrium point can be any unspecified convex combination of the local equilibria of all the agents in the absence of communication. Both the cases with constant and time-varying step-sizes are considered. In the case when the convex combination is required to be a straight average and interaction between any pair of neighboring agents may be uni-directional, so that doubly stochastic matrices cannot be implemented in a distributed manner, the paper proposes a push-sum-type distributed stochastic approximation algorithm and provides its finite-time bound for the time-varying step-size case by leveraging the analysis for the consensus-type algorithm with stochastic matrices and developing novel properties of the push-sum algorithm. Distributed temporal difference learning is discussed as an illustrative application.
A shortened version has been accepted by Automatica
References in corpus (9)
- Fully Decentralized Multi-Agent Reinforcement Learning with Networked Agents
- Finite-Sample Analysis for SARSA with Linear Function Approximation
- Finite-Time Analysis of Distributed TD(0) with Linear Function Approximation for Multi-Agent Reinforcement Learning
- Two Time-scale Off-Policy TD Learning: Non-asymptotic Analysis over Markovian Samples
- Finite Sample Analysis of Two-Timescale Stochastic Approximation with Applications to Reinforcement Learning
- A Multistep Lyapunov Approach for Finite-Time Analysis of Biased Stochastic Approximation
- Finite Sample Analysis of the GTD Policy Evaluation Algorithms in Markov Setting
- Finite-Time Performance Bounds and Adaptive Learning Rate Selection for Two Time-Scale Reinforcement Learning
- Greedy-GQ with Variance Reduction: Finite-time Analysis and Improved Complexity