1 paper
Enhan Li, Junhao He, Hongyang Du
On-policy distillation (OPD) supervises a student language model on trajectories sampled from its current policy, but assigns equal credit to response tokens with unequal supervisi…