1 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Shipeng Li, Zhiqin Yang, Shikun Li +7
Reinforcement learning with verifiable rewards (RLVR) has become a key technique for enhancing LLMs' reasoning abilities, yet its data inefficiency remains a major bottleneck. To a…