7 citations · 11 across the 13 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2023
Partial advantage estimator for proximal policy optimization
Xiulei Song, Yizhao Jin, Greg Slabaugh +1
Estimation of value in policy gradient methods is a fundamental problem. Generalized Advantage Estimation (GAE) is an exponentially-weighted estimator of an advantage function simi…
cs.LG2023★ 1 cited
Joint action loss for proximal policy optimization
Xiulei Song, Yizhao Jin, Greg Slabaugh +1
PPO (Proximal Policy Optimization) is a state-of-the-art policy gradient algorithm that has been successfully applied to complex computer games such as Dota 2 and Honor of Kings. I…