1 paper · 1 filter
Ke Sun, Yizhou Zhao, Jiayi Xin +2
Context or prompt-level reweighting has emerged as a central algorithmic lever in Reinforcement Learning with Verified Rewards (RLVR) for improving the reasoning capability of larg…