1 paper · 1 filter
Bingqing Jiang, Difan Zou
Reinforcement learning from verifiable rewards (RLVR), especially with Group Relative Policy Optimization (GRPO), has shown strong potential for improving the reasoning capabilitie…