15 citations · 17 across the 18 of their papers we have counts for
1 paper · 2 filters
Yuhang Li, Reena Elangovan, Xin Dong +2
Reinforcement learning with verifiable rewards (RLVR) has become a trending paradigm for training reasoning large language models (LLMs). However, due to the autoregressive decodin…