1 paper
Anton Baumann, Akmal Ashirmatov, Leo Schmidt-Traub +4
On-policy self-distillation provides dense, token-level supervision by conditioning a model on privileged information and distilling the resulting teacher distribution back into th…