2 papers
cs.LG2025
Mitigating Think-Answer Mismatch in LLM Reasoning Through Noise-Aware Advantage Reweighting
Si Shen, Peijun Shen, Wenhua Zhao +1
Group-Relative Policy Optimization (GRPO) is a key technique for training large reasoning models, yet it suffers from a critical vulnerability: the \emph{Think-Answer Mismatch}, wh…
math.AG2024
A local to global question for linear functionals
George F. Seelinger, Wenhua Zhao
Let be an algebraically closed field and let . Consider with standard basis and its dual space $V^*= {\mathrm{Hom}}_{F-{\mathr…