1 paper · 1 filter
Sihan Wang, Xiyao Liu, Lianqing Liu +1
On-policy self-distillation (OPSD) trains a model on its own rollouts and uses a frozen copy to provide dense token-level targets conditioned on a reference target. This works well…