Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Short Chains, Deep Thoughts: Balancing Reasoning Efficiency and Intra-Segment Capability via Split-Merge Optimization
Runquan Gui, Jie Wang, Zhihai Wang +3
While Large Reasoning Models (LRMs) have demonstrated impressive capabilities in solving complex tasks through the generation of long reasoning chains, this reliance on verbose gen…
cs.CL2026
FaithRL: Learning to Reason Faithfully through Step-Level Faithfulness Maximization
Runquan Gui, Yafu Li, Xiaoye Qu +3
Reinforcement Learning with Verifiable Rewards (RLVR) has markedly improved the performance of Large Language Models (LLMs) on tasks requiring multi-step reasoning. However, most R…