2 papers
cs.CL2026
FABSVer: Faster Training and Better Self-Verification for LLM Mathematical Reasoning
Haihui Pan, Junwei Bao, Hongfei Jiang +1
While large language models have made significant progress in mathematical reasoning, they remain unreliable at judging the correctness of their own solutions. Existing approaches…
cs.CL2026
Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity
Haihui Pan, Yuzhong Hong, Kaichen Zhang +4
In many large language model (LLM) alignment applications, users expect not only high-quality outputs but also substantial diversity. However, existing methods often face a fundame…