2 citations · 2 across the 1 of their papers we have counts for
1 paper · 1 filter
Jun Song, Chaoyue Zhao
Trust Region Policy Optimization (TRPO) and Proximal Policy Optimization (PPO), as the widely employed policy based reinforcement learning (RL) methods, are prone to converge to a…