21 citations · 43 across the 3 of their papers we have counts for
3 papers
Multi-step Off-policy Learning Without Importance Sampling Ratios
Ashique Rupam Mahmood, Huizhen Yu, Richard S. Sutton
To estimate the value functions of policies from exploratory data, most model-free off-policy algorithms rely on importance sampling, where the use of importance sampling ratios of…
Emphatic Temporal-Difference Learning
A. Rupam Mahmood, Huizhen Yu, Martha White +1
Emphatic algorithms are temporal-difference learning algorithms that change their effective state distribution by selectively emphasizing and de-emphasizing their updates on differ…
An Empirical Evaluation of True Online TD(λ)
Harm van Seijen, A. Rupam Mahmood, Patrick M. Pilarski +1
The true online TD(λ) algorithm has recently been proposed (van Seijen and Sutton, 2014) as a universal replacement for the popular TD(λ) algorithm, in temporal-difference learning…