Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Multi-Turn On-Policy Distillation with Prefix Replay
Baohao Liao, Hanze Dong, Christof Monz +3
We study on-policy distillation (OPD) for agentic tasks, where an LLM agent interacts with an environment over multiple turns and a student imitates a teacher over these multi-turn…
cs.LG2025
Fractured Chain-of-Thought Reasoning
Baohao Liao, Hanze Dong, Yuhui Xu +4
Inference-time scaling techniques have significantly bolstered the reasoning capabilities of large language models (LLMs) by harnessing additional computational effort at inference…
cs.LG2025
Reward Models Identify Consistency, Not Causality
Yuhui Xu, Hanze Dong, Lei Wang +2
Reward models (RMs) play a crucial role in aligning large language models (LLMs) with human preferences and enhancing reasoning quality. Traditionally, RMs are trained to rank cand…