1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
Si Shen, Peijun Shen, Wenhua Zhao +1
Group-Relative Policy Optimization (GRPO) is a key technique for training large reasoning models, yet it suffers from a critical vulnerability: the \emph{Think-Answer Mismatch}, wh…