1 paper · 1 filter
Liang Chen, Xueting Han, Qizhou Wang +4
Balancing exploration and exploitation remains a central challenge in reinforcement learning with verifiable rewards (RLVR) for large language models (LLMs). Current RLVR methods o…