275 citations · 302 across the 21 of their papers we have counts for
1 paper · 2 filters
Gal Dalal, Balazs Szorenyi, Gugan Thoppe
Policy evaluation in reinforcement learning is often conducted using two-timescale stochastic approximation, which results in various gradient temporal difference methods such as G…