15 citations · 24 across the 15 of their papers we have counts for
1 paper · 1 filter
Jiamian Wang, Samyadeep Basu, Koustava Goswami +2
Group Relative Policy Optimization (GRPO) is widely studied for reinforcement learning with verifiable rewards, where its advantage estimator assigns each rollout a magnitude from…