1 paper
Perttu Hämäläinen, Amin Babadi, Xiaoxiao Ma +1
Proximal Policy Optimization (PPO) is a highly popular model-free reinforcement learning (RL) approach. However, we observe that in a continuous action space, PPO can prematurely s…