1 paper · 1 filter
Xiaolong Jin, Xuandong Zhao, Wenbo Guo +2
Reinforcement learning with verifiable rewards (RLVR) has driven substantial progress in large language model reasoning, but relies on ground-truth supervision that is costly or in…