1 paper
Yaozhong Gan, Renye Yan, Xiaoyang Tan +2
Proximal Policy Optimization (PPO) is a popular model-free reinforcement learning algorithm, esteemed for its simplicity and efficacy. However, due to its inherent on-policy nature…