3 papers
cs.LG2026
A Token-Level Analysis of Sampled-Token Reverse-KL On-Policy Distillation
Bing Shao, Jiazheng Zhang, Long Ma +11
On-policy distillation (OPD) supervises a student on its own trajectories with token-level signals from a frozen teacher, yet how a sampled loss allocates updates across tokens rem…
cs.AI2026
Thinking with Constructions: A Benchmark and Policy Optimization for Visual-Text Interleaved Geometric Reasoning
Haokun Zhao, Wanshi Xu, Haidong Yuan +3
Geometric reasoning inherently requires "thinking with constructions" -- the dynamic manipulation of visual aids to bridge the gap between problem conditions and solutions. However…
cs.AI2026
Steering LLMs via Scalable Interactive Oversight
Enyu Zhou, Zhiheng Xi, Long Ma +9
As Large Language Models increasingly automate complex, long-horizon tasks such as \emph{vibe coding}, a supervision gap has emerged. While models excel at execution, users often s…