1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Nil Stolt Ansó
The Sampled Policy Gradient (SPG) algorithm is a new offline actor-critic variant that samples in the action space to approximate the policy gradient. It does so by using the criti…