1 paper
Zhuofan Chen, Ziqian Jiao, Yikai Cui +3
Reinforcement learning with verifiable rewards (RLVR) is sensitive to which problems a model trains on, yet existing selection criteria--difficulty filtering, hand-curation, reward…