1 paper
Zhizhao Liu, Zhiliang Tian, Xi Wang +4
Reinforcement learning with verifiable rewards (RLVR) improves the reasoning capabilities of large language models but relies on costly rollout exploration. Assigning the same expl…