Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
: A Generalist Value Model for Any Policy at State Zero
Yi-Kai Zhang, Zhiyuan Yao, Hongyan Hao +6
Policy gradient methods rely on a baseline to measure the relative advantage of an action, ensuring the model reinforces behaviors that outperform its current average capability. I…
cs.CL2026
Learning to Self-Verify Makes Language Models Better Reasoners
Yuxin Chen, Yu Wang, Yi Zhang +9
Recent large language models (LLMs) achieve strong performance in generating promising reasoning paths for complex tasks. However, despite powerful generation ability, LLMs remain…