15 citations · 15 across the 1 of their papers we have counts for
1 paper
Vincent Mai, Kaustubh Mani, Liam Paull
In model-free deep reinforcement learning (RL) algorithms, using noisy value estimates to supervise policy evaluation and optimization is detrimental to the sample efficiency. As t…