7 citations · 7 across the 6 of their papers we have counts for
1 paper · 1 filter
Benjamin Turtel, Danny Franklin, Kris Skotheim +2
Reinforcement Learning with Verifiable Rewards (RLVR) has been an effective approach for improving Large Language Models' reasoning in domains such as coding and mathematics. Here,…