1 paper · 1 filter
Konstantin Dobler, Federico Scozzafava, Jonathan Janke +2
Reinforcement Learning with Verifiable Rewards (RLVR), often optimized with Group Relative Policy Optimization (GRPO), has become a central recipe for improving the reasoning capab…