1 citations · 1 across the 4 of their papers we have counts for
Showing 2026Show all
3 papers · 1 filter
cs.CL2026
FABSVer: Faster Training and Better Self-Verification for LLM Mathematical Reasoning
Haihui Pan, Junwei Bao, Hongfei Jiang +1
While large language models have made significant progress in mathematical reasoning, they remain unreliable at judging the correctness of their own solutions. Existing approaches…
cs.CL2026
Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity
Haihui Pan, Yuzhong Hong, Kaichen Zhang +4
In many large language model (LLM) alignment applications, users expect not only high-quality outputs but also substantial diversity. However, existing methods often face a fundame…
cs.CL2026
Elo-Evolve: A Co-evolutionary Framework for Language Model Alignment
Jing Zhao, Ting Zhen, Junwei Bao +2
Current alignment methods for Large Language Models (LLMs) rely on compressing vast amounts of human preference data into static, absolute reward functions, leading to data scarcit…