50 citations · 136 across the 15 of their papers we have counts for
1 paper · 1 filter
Zhongxi Qiu, Zhang Zhang, Yan Hu +2
This paper explores optimal data selection strategies for Reinforcement Learning with Verified Rewards (RLVR) training in the medical domain. While RLVR has shown exceptional poten…