2 citations · 2 across the 1 of their papers we have counts for
1 paper
Jun Song, Chaoyue Zhao
Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO), as the widely employed policy based reinforcement learning (RL) methods, are prone to converge to a…