1 paper
Yuning Wu, Ke Wang, Devin Chen +1
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising paradigm for post-training reasoning models. However, group-based methods such as Group Relative Po…