1 paper
Zhongwei Wan, Yun Shen, Zhihao Dou +9
Reinforcement learning with verifiers (RLVR) is a central paradigm for improving large language model (LLM) reasoning, yet existing methods often suffer from limited exploration. P…