1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Ben Rahman
Despite Proximal Policy Optimization (PPO) dominating policy gradient methods -- from robotic control to game AI -- its static trust region forces a brittle trade-off: aggressive c…