1 paper · 1 filter
Hwanwoo Kim, Dongkyu Derek Cho, Eric Laber
Temporal difference (TD) learning is a cornerstone of reinforcement learning. In the average-reward setting, standard TD(I^») is highly sensitive to the choice of step-size and th…