1 paper
Zhongxi Qiu, Zhang Zhang, Yan Hu +2
This paper explores optimal data selection strategies for Reinforcement Learning with Verified Rewards (RLVR) training in the medical domain. While RLVR has shown exceptional poten…