1 paper
Jiaxin Guo, Yanwei Yue, Xuanbo Fan +2
On-policy self-distillation improves language-model reasoning by querying a teacher on states actually visited by the student. Recent methods create a powerful information asymmetr…