1 paper
Ryo Mitsuhashi, Patrick Chen, Isabelle Tseng +2
Reinforcement learning with verifiable rewards (RLVR) has become a standard paradigm for post-training language models, but in practice, verifiers are rarely perfect. Recent theore…