1 paper
Minwu Kim, Safal Shrestha, Anubhav Shrestha +1
As Reinforcement Learning with Verifiable Rewards (RLVR) substantially improves the reasoning abilities of large language models (LLMs), a new bottleneck emerges: more training pro…