3 citations · 3 across the 2 of their papers we have counts for
1 paper · 1 filter
Jayesh Singla, Ananye Agarwal, Deepak Pathak
Despite extreme sample inefficiency, on-policy reinforcement learning, aka policy gradients, has become a fundamental tool in decision-making problems. With the recent advances in…