1 paper
Shenzhi Yang, Guangcheng Zhu, Bowen Song +7
Reinforcement Learning with Verifiable Rewards (RLVR) effectively trains reasoning models that rely on abundant perfect labels, but its vulnerability to unavoidable noisy labels du…