10 citations · 10 across the 6 of their papers we have counts for
1 paper · 1 filter
Razvan-Andrei Lascu, David Šiška, Łukasz Szpruch
Proximal Policy Optimization (PPO) is widely used in reinforcement learning due to its strong empirical performance, yet it lacks formal guarantees for policy improvement and conve…