2 citations · 2 across the 2 of their papers we have counts for
1 paper · 1 filter
Yunke Ao, Le Chen, Bruce D. Lee +5
Proximal Policy Optimization (PPO) has become the predominant algorithm for on-policy reinforcement learning due to its scalability and empirical robustness across domains. However…