12 papers
Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning
Xinyan Guan, Jiali Zeng, Chunlei Xin +5
Large language models generate computationally expensive yet semantically void reasoning on beyond-capability tasks, creating risks where plausible-sounding but incorrect derivatio…
Offline Exploration-Aware Fine-Tuning for Long-Chain Mathematical Reasoning
Yongyu Mu, Jiali Zeng, Fandong Meng +2
Through encouraging self-exploration, reinforcement learning from verifiable rewards (RLVR) has significantly advanced the mathematical reasoning capabilities of large language mod…
GRAM-R: Self-Training Generative Foundation Reward Models for Reward Reasoning
Chenglong Wang, Yongyu Mu, Hang Zhou +10
Significant progress in reward modeling over recent years has been driven by a paradigm shift from task-specific designs towards generalist reward models. Despite this trend, devel…
Dissecting Long-Chain-of-Thought Reasoning Models: An Empirical Study
Yongyu Mu, Jiali Zeng, Bei Li +5
Despite recent progress in training long-chain-of-thought reasoning models via scaling reinforcement learning (RL), its underlying training dynamics remain poorly understood, and s…
ConCISE: Confidence-guided Compression in Step-by-step Efficient Reasoning
Ziqing Qiao, Yongheng Deng, Jiali Zeng +7
Large Reasoning Models (LRMs) perform strongly in complex reasoning tasks via Chain-of-Thought (CoT) prompting, but often suffer from verbose outputs, increasing computational over…
RewardAnything: Generalizable Principle-Following Reward Models
Zhuohao Yu, Jiali Zeng, Weizheng Gu +7
Reward Models, essential for guiding Large Language Model optimization, are typically trained on fixed preference datasets, resulting in rigid alignment to single, implicit prefere…