1 paper
Junzi Zhang, Jongho Kim, Brendan O'Donoghue +1
Policy gradient methods are among the most effective methods for large-scale reinforcement learning, and their empirical success has prompted several works that develop the foundat…