1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.LG2025
Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting
Si Shen, Peijun Shen, Wenhua Zhao +1
Group-Relative Policy Optimization (GRPO) is a key technique for training large reasoning models, yet it suffers from a critical vulnerability: the \emph{Think-Answer Mismatch}, wh…
cs.CL2025★ 1 cited
Long Is More Important Than Difficult for Training Reasoning Models
Si Shen, Fei Huang, Zhixiao Zhao +3
Difficult problems, which often result in long reasoning traces, are widely recognized as key factors for enhancing the performance of reasoning models. However, such high-challeng…