3 citations · 3 across the 1 of their papers we have counts for
1 paper
Yuhua Zhu, Zach Izzo, Lexing Ying
In model-free reinforcement learning, the temporal difference method and its variants become unstable when combined with nonlinear function approximations. Bellman residual minimiz…