1 paper · 1 filter
Mingxiong Lin, Zhangquan Gong, Maowen Tang +6
Reinforcement Learning with Verifiable Rewards (RLVR) has become the standard paradigm for LLM mathematical reasoning, where Group Relative Policy Optimization (GRPO) serves as the…