1 paper
Feng Zhang, Xinhong Ma, Ziqiang Dong +5
Group Relative Policy Optimization (GRPO) is one of the most widely adopted RLVR algorithms for post-training large language models on reasoning tasks. We first show that GRPO admi…