works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.LG2026

Multi-Turn On-Policy Distillation with Prefix Replay

Baohao Liao, Hanze Dong, Christof Monz +3

The paper introduces ReOPD, a method that reuses pre‑collected teacher trajectories as replayed prefixes to train LLM agents without costly new environment interactions, improving…

cs.LG2026

Fractured Chain-of-Thought Reasoning

Baohao Liao, Hanze Dong, Yuhui Xu +4

Inference-time scaling techniques have significantly bolstered the reasoning capabilities of large language models (LLMs) by harnessing additional computational effort at inference…

cs.CL2026

LiveMathematicianBench: A Live Benchmark for Mathematician-Level Reasoning with Proof Sketches

Linyang He, Qiyao Yu, Hanze Dong +5

Mathematical reasoning is a hallmark of human intelligence, and whether large language models (LLMs) can meaningfully perform it remains a central question in artificial intelligen…

cs.CL2025

Beyond 'Aha!': Toward Systematic Meta-Abilities Alignment in Large Reasoning Models

Zhiyuan Hu, Yibo Wang, Hanze Dong +5

Large reasoning models (LRMs) already possess a latent capacity for long chain-of-thought reasoning. Prior work has shown that outcome-based reinforcement learning (RL) can inciden…

cs.LG2025

Reward Models Identify Consistency, Not Causality

Yuhui Xu, Hanze Dong, Lei Wang +2

Reward models (RMs) play a crucial role in aligning large language models (LLMs) with human preferences and enhancing reasoning quality. Traditionally, RMs are trained to rank cand…