1 paper · 1 filter
Bing Shao, Jiazheng Zhang, Long Ma +11
On-policy distillation (OPD) supervises a student on its own trajectories with token-level signals from a frozen teacher, yet how a sampled loss allocates updates across tokens rem…