1 paper
Haotian Wang, Lian Yan, Xingzhi Yao +4
In Reinforcement Learning with Verifiable Rewards (RLVR) frameworks for mathematical reasoning tasks, floating-point results are typically evaluated using a tolerance-based reward.…