2 papers
cs.LG2026
ReOrder-OPD:Reliability-Aware Prompt Ordering for On-Policy Distillation
Ximo Zhu, Ruiqi Liu, Rong Wang +8
On-policy distillation (OPD) applies token-level teacher supervision to student-generated trajectories, but this supervision is not always reliable. Existing methods use local conf…
cs.CV2026
Visual-Advantage On-Policy Distillation for Vision-Language Models
Ruiqi Liu, Xiaolei Lv, Gengsheng Li +8
On-policy knowledge distillation has proven effective for language models, yet its application to vision-language models (VLMs) remains underexplored. We observe that standard on-p…