33 citations · 50 across the 2 of their papers we have counts for
2 papers
cs.LG2020★ 33 cited
Non-asymptotic Convergence Analysis of Two Time-scale (Natural) Actor-Critic Algorithms
Tengyu Xu, Zhe Wang, Yingbin Liang
As an important type of reinforcement learning algorithms, actor-critic (AC) and natural actor-critic (NAC) algorithms are often executed in two ways for finding optimal policies.…
cs.LG2020★ 17 cited
Reanalysis of Variance Reduced Temporal Difference Learning
Tengyu Xu, Zhe Wang, Yi Zhou +1
Temporal difference (TD) learning is a popular algorithm for policy evaluation in reinforcement learning, but the vanilla TD can substantially suffer from the inherent optimization…