Finite-Time Analysis of Asynchronous Stochastic Approximation and -Learning
arXiv:2002.00260
Abstract
We consider a general asynchronous Stochastic Approximation (SA) scheme featuring a weighted infinity-norm contractive operator, and prove a bound on its finite-time convergence rate on a single trajectory. Additionally, we specialize the result to asynchronous -learning. The resulting bound matches the sharpest available bound for synchronous -learning, and improves over previous known bounds for asynchronous -learning.
References in corpus (4)
- Two Time-scale Off-Policy TD Learning: Non-asymptotic Analysis over Markovian Samples
- Finite-Time Error Bounds For Linear Stochastic Approximation and TD Learning
- Stochastic approximation with cone-contractive operators: Sharp -bounds for -learning
- Model-free Reinforcement Learning in Infinite-horizon Average-reward Markov Decision Processes