1 paper
Andy Xu, Rohan Desai, Larry Wang +2
Reinforcement Learning from Verifiable Rewards (RLVR) has emerged as a promising approach to improve correctness in LLMs, however, in many scientific problems, the objective is not…