5 citations · 14 across the 21 of their papers we have counts for
1 paper · 2 filters
Xitai Jiang, Zihan Tang, Wenze Lin +3
Reinforcement learning from verifiable rewards (RLVR) has shown strong promise for LLM reasoning, but outcome-based RLVR remains inefficient on hard problems because correct final-…