Finite Sample Analysis of Two-Timescale Stochastic Approximation with Applications to Reinforcement Learning
arXiv:1703.05376
Abstract
Two-timescale Stochastic Approximation (SA) algorithms are widely used in Reinforcement Learning (RL). Their iterates have two parts that are updated using distinct stepsizes. In this work, we develop a novel recipe for their finite sample analysis. Using this, we provide a concentration bound, which is the first such result for a two-timescale SA. The type of bound we obtain is known as `lock-in probability'. We also introduce a new projection scheme, in which the time between successive projections increases exponentially. This scheme allows one to elegantly transform a lock-in probability into a convergence rate result for projected two-timescale SA. From this latter result, we then extract key insights on stepsize selection. As an application, we finally obtain convergence rates for the projected two-timescale RL algorithms GTD(0), GTD2, and TDC.
Cited by in corpus (14)
- Non-asymptotic Convergence Analysis of Two Time-scale (Natural) Actor-Critic Algorithms
- Reanalysis of Variance Reduced Temporal Difference Learning
- A Tale of Two-Timescale Reinforcement Learning with the Tightest Finite-Time Bound
- Calibration of Shared Equilibria in General Sum Partially Observable Markov Games
- Sample Complexity Bounds for Two Timescale Value-based Reinforcement Learning Algorithms
- Finite-Time Convergence Rates of Nonlinear Two-Time-Scale Stochastic Approximation under Markovian Noise
- Finite-Time Error Bounds for Distributed Linear Stochastic Approximation
- Multi-Agent Off-Policy TD Learning: Finite-Time Analysis with Near-Optimal Sample Complexity and Communication Complexity
- Distributed TD(0) with Almost No Communication
- PER-ETD: A Polynomially Efficient Emphatic Temporal Difference Learning Method
- Random directions stochastic approximation with deterministic perturbations
- Dynamics of stochastic approximation with iterate-dependent Markov noise under verifiable conditions in compact state space with the stability of iterates not ensured
- On Convergence of Gradient Expected Sarsa()
- Expected Sarsa() with Control Variate for Variance Reduction