13 citations · 37 across the 31 of their papers we have counts for
1 paper · 1 filter
Jiaming Guo, Rui Zhang, Xishan Zhang +6
Policy gradient methods are appealing in deep reinforcement learning but suffer from high variance of gradient estimate. To reduce the variance, the state value function is applied…