1 paper
Vaskar Nath, Elaine Lau, Anisha Gunjal +3
We study the process through which reasoning models trained with reinforcement learning on verifiable rewards (RLVR) can learn to solve new problems. We find that RLVR drives perfo…