1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Yuze Gao
Reinforcement learning from verifiable rewards (RLVR) improves reasoning even when the reward signal is spurious -- assigning credit to the group-plurality answer rather than a gro…