1 paper
Samyak Shrestha, Alexander Tessier
On-policy self-distillation (OPSD) trains a student on its own responses using token-level supervision from the same model conditioned on privileged reference information. We inves…