16 citations · 16 across the 2 of their papers we have counts for
1 paper · 1 filter
Saurabh Garg, Joshua Zhanson, Emilio Parisotto +6
Modern policy gradient algorithms such as Proximal Policy Optimization (PPO) rely on an arsenal of heuristics, including loss clipping and gradient clipping, to ensure successful l…