10 citations · 10 across the 1 of their papers we have counts for
1 paper
Bowen Zheng, Ran Cheng
While off-policy reinforcement learning (RL) algorithms are sample efficient due to gradient-based updates and data reuse in the replay buffer, they struggle with convergence to lo…