147 citations · 254 across the 59 of their papers we have counts for
1 paper · 2 filters
Zhihang Lin, Mingbao Lin, Yuan Xie +1
This paper introduces Completion Pruning Policy Optimization (CPPO) to accelerate the training of reasoning models based on Group Relative Policy Optimization (GRPO). GRPO, while e…