1 paper
Yunho Choi, Jongwon Lim, Woojin Ahn +3
Reinforcement learning with verifiable rewards (RLVR) for Large Reasoning Models hinges on baseline estimation for variance reduction, but existing approaches pay a heavy price: PP…