1 paper
Zihao Han, Tiangang Zhang, Huaibin Wang +1
On-policy self-distillation has become a strong recipe for LLM reasoning, where a privileged teacher supervises the student's own rollouts while conditioning on the reference solut…