1 paper
Gaurisankar Jayadas, Aske Plaat, Álvaro Serra-Gómez +1
Reinforcement Learning with Verifiable Rewards (RLVR) performs well on problems with clear rewards, such as mathematics and coding, but whether it also works where the reward is le…