1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Xinyu Zhu, Mengzhou Xia, Zhepei Wei +3
Reinforcement learning with verifiable rewards (RLVR) is a promising approach for training language models (LMs) on reasoning tasks that elicit emergent long chains of thought (CoT…