1 paper
Devan Shah, Owen Yang, Daniel Yang +2
Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning abilities of large language models (LLMs) on mathematics and programming tasks, but standard approa…