11 citations · 23 across the 5 of their papers we have counts for
1 paper · 1 filter
Anton Orell Wiehe, Nil Stolt Ansó, Madalina M. Drugan +1
In this paper, a new offline actor-critic learning algorithm is introduced: Sampled Policy Gradient (SPG). SPG samples in the action space to calculate an approximated policy gradi…