1 paper · 1 filter
Zelal Su, Mustafaoglu, Sungyoung Lee +3
Proximal policy optimization (PPO) approximates the trust region update using multiple epochs of clipped SGD. Each epoch may drift further from the natural gradient direction, crea…