1 paper
Jie Zhang, Jingxiao Yang, Zhehao Huang +2
Reinforcement learning with verifiable rewards (RLVR) supervises mathematical reasoning through final-answer correctness, but provides little guidance on individual tokens. On-poli…