1 paper
Yuchun Miao, Sen Zhang, Yuqi Zhang +4
Reinforcement Learning with Verifiable Rewards (RLVR) has become the dominant paradigm for advanced reasoning in Large Language Models (LLMs), but rollout samples are expensive to…