Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Mismatch Matters: On-Policy Distillation Beyond Token Agreement
Zichao Yu, Chengzhi Yu, Shengze Xu +4
On-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerate agreement, where students exploit repet…
cs.AI2026
SCULPT: Constraint-Guided Pruned MCTS that Carves Efficient Paths for Mathematical Reasoning
Qitong Fang, Haotian Li, Xu Wang
Automated agent workflows can enhance the problem-solving ability of large language models (LLMs), but common search strategies rely on stochastic exploration and often traverse im…