4 papers
Beyond variance reduction: Understanding the true impact of baselines on policy optimization
Wesley Chung, Valentin Thomas, Marlos C. Machado +1
Bandit and reinforcement learning (RL) problems can often be framed as optimization problems where the goal is to maximize average performance while having access only to stochasti…
Incrementally Learning Functions of the Return
Brendan Bennett, Wesley Chung, Muhammad Zaheer +1
Temporal difference methods enable efficient estimation of value functions in reinforcement learning in an incremental fashion, and are of broader interest because they correspond…
Importance Resampling for Off-policy Prediction
Matthew Schlegel, Wesley Chung, Daniel Graves +2
Importance sampling (IS) is a common reweighting strategy for off-policy prediction in reinforcement learning. While it is consistent and unbiased, it can result in high variance u…
High-confidence error estimates for learned value functions
Touqir Sajed, Wesley Chung, Martha White
Estimating the value function for a fixed policy is a fundamental problem in reinforcement learning. Policy evaluation algorithms---to estimate value functions---continue to be dev…