From the 1 of 5 linked papers with an AI index.
5 papers
Multi-Turn On-Policy Distillation with Prefix Replay
Baohao Liao, Hanze Dong, Christof Monz +3
The paper introduces ReOPD, a method that reuses pre‑collected teacher trajectories as replayed prefixes to train LLM agents without costly new environment interactions, improving…
Fractured Chain-of-Thought Reasoning
Baohao Liao, Hanze Dong, Yuhui Xu +4
Inference-time scaling techniques have significantly bolstered the reasoning capabilities of large language models (LLMs) by harnessing additional computational effort at inference…
LiveMathematicianBench: A Live Benchmark for Mathematician-Level Reasoning with Proof Sketches
Linyang He, Qiyao Yu, Hanze Dong +5
Mathematical reasoning is a hallmark of human intelligence, and whether large language models (LLMs) can meaningfully perform it remains a central question in artificial intelligen…
Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models
Zhiyuan Hu, Yibo Wang, Hanze Dong +5
Large reasoning models (LRMs) already possess a latent capacity for long chain-of-thought reasoning. Prior work has shown that outcome-based reinforcement learning (RL) can inciden…
Reward Models Identify Consistency, Not Causality
Yuhui Xu, Hanze Dong, Lei Wang +2
Reward models (RMs) play a crucial role in aligning large language models (LLMs) with human preferences and enhancing reasoning quality. Traditionally, RMs are trained to rank cand…