2 papers
cs.CL2026
FaithRL: Learning to Reason Faithfully through Step-Level Faithfulness Maximization
Runquan Gui, Yafu Li, Xiaoye Qu +3
Reinforcement Learning with Verifiable Rewards (RLVR) has markedly improved the performance of Large Language Models (LLMs) on tasks requiring multi-step reasoning. However, most R…
cs.AI2025
HyperTree Planning: Enhancing LLM Reasoning via Hierarchical Thinking
Runquan Gui, Zhihai Wang, Jie Wang +7
Recent advancements have significantly enhanced the performance of large language models (LLMs) in tackling complex reasoning tasks, achieving notable success in domains like mathe…