3 citations · 3 across the 3 of their papers we have counts for
1 paper · 1 filter
Nonghai Zhang, Weitao Ma, Zhanyu Ma +5
Group Relative Policy Optimization (GRPO) significantly enhances the reasoning performance of Large Language Models (LLMs). However, this success heavily relies on expensive extern…