1 paper · 1 filter
Junhua Chen, Zixi Zhang, Hantao Zhong +1
We introduce Group Policy Gradient (GPG), a family of critic-free policy-gradient estimators for general MDPs. Inspired by the success of GRPO's approach in Reinforcement Learning…