3 papers
cs.CL2026
ReGuLaR: Variational Latent Reasoning Guided by Rendered Chain-of-Thought
Fanmeng Wang, Haotian Liu, Guojiang Zhao +2
While Chain-of-Thought (CoT) significantly enhances the performance of Large Language Models (LLMs), explicit reasoning chains introduce substantial computational redundancy. Recen…
cs.LG2025
CGSPG: Confidence-calibrated Group Sequence Policy Gradient towards Self-aware Reasoning
Haotian Liu, Shuo Wang, Hongteng Xu
Reinforcement Learning (RL) methods, exemplified by Group Relative Policy Optimization (GRPO) and its variants, play a central role in developing reasoning models. However, these m…
cs.CL2025
Learning an Effective Premise Retrieval Model for Efficient Mathematical Formalization
Yicheng Tao, Haotian Liu, Shanwen Wang +1
Formalized mathematics has recently garnered significant attention for its ability to assist mathematicians across various fields. Premise retrieval, as a common step in mathematic…