1 paper
Song Yu, Li Li, Wenwen Zhao +1
Reinforcement learning with verifiable rewards (RLVR), particularly Group Relative Policy Optimization (GRPO), has advanced LLM reasoning. However, GRPO suffers from three credit a…