1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Xiaoyang Yuan, Yujuan Ding, Yi Bin +5
Reinforcement Learning with Verifiable Rewards (RLVR) is a promising paradigm for enhancing the reasoning ability in Large Language Models (LLMs). However, prevailing methods prima…