3 papers
cs.CL2026
EasyOPD: An Easy-to-use On-Policy Distillation Framework for Large Language Models
Jie Sun, Mao Zheng, Mingyang Song +7
Conventional language-model distillation often relies on fixed teacher-generated data, which may not cover the states encountered by an evolving student policy. On-policy distillat…
cs.AI2026
TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety
Zhepei Hong, Lin Wang, Liting Li +5
Long-horizon LLM agents produce safety evidence across long trajectories, where sparse, delayed, and compositional risk signals often escape local moderation. Existing turn-level o…
cs.LG2026
Rubric-based On-policy Distillation
Junfeng Fang, Zhepei Hong, Mao Zheng +7
On-policy distillation (OPD) is a powerful paradigm for model alignment, yet its reliance on teacher logits restricts its application to white-box scenarios. We contend that struct…