1 paper · 1 filter
Yuzhong Zhao, Yue Liu, Junpeng Liu +9
Group Relative Policy Optimization (GRPO) has significantly enhanced the reasoning capability of large language models by optimizing the arithmetic mean of token-level rewards. Unf…