1 paper · 2 filters
Craig Sherstan, Brendan Bennett, Kenny Young +4
This paper investigates estimating the variance of a temporal-difference learning agent's update target. Most reinforcement learning methods use an estimate of the value function,…