4 citations · 6 across the 2 of their papers we have counts for
3 papers
Target Network and Truncation Overcome The Deadly Triad in -Learning
Zaiwei Chen, John Paul Clarke, Siva Theja Maguluri
-learning with function approximation is one of the most empirically successful while theoretically mysterious reinforcement learning (RL) algorithms, and was identified in Sutt…
Finite-Sample Analysis of Off-Policy TD-Learning via Generalized Bellman Operators
Zaiwei Chen, Siva Theja Maguluri, Sanjay Shakkottai +1
In temporal difference (TD) learning, off-policy sampling is known to be more practical than on-policy sampling, and by decoupling learning from data collection, it enables data re…
Finite-Sample Analysis of Off-Policy Natural Actor-Critic Algorithm
Sajad Khodadadian, Zaiwei Chen, Siva Theja Maguluri
In this paper, we provide finite-sample convergence guarantees for an off-policy variant of the natural actor-critic (NAC) algorithm based on Importance Sampling. In particular, we…