1 citations · 1 across the 6 of their papers we have counts for
1 paper · 1 filter
Kangda Wei, Ruihong Huang
Group Relative Policy Optimization (GRPO) has become a standard approach for training mathematical reasoning models; however, its reliance on multiple completions per prompt makes…