21 citations · 49 across the 6 of their papers we have counts for
Showing 2017Show all
3 papers · 1 filter
cs.LG2017
On Convergence of some Gradient-based Temporal-Differences Algorithms for Off-Policy Learning
Huizhen Yu
We consider off-policy temporal-difference (TD) learning methods for policy evaluation in Markov decision processes with finite spaces and discounted reward criteria, and we presen…
cs.LG2017
On Generalized Bellman Equations and Temporal-Difference Learning
Huizhen Yu, A. Rupam Mahmood, Richard S. Sutton
We consider off-policy temporal-difference (TD) learning in discounted Markov decision processes, where the goal is to evaluate a policy in a model-free way by using observations o…
cs.LG2017★ 21 cited
Multi-step Off-policy Learning Without Importance Sampling Ratios
Ashique Rupam Mahmood, Huizhen Yu, Richard S. Sutton
To estimate the value functions of policies from exploratory data, most model-free off-policy algorithms rely on importance sampling, where the use of importance sampling ratios of…