1 paper
Junfeng Fang, Zhepei Hong, Mao Zheng +7
On-policy distillation (OPD) is a powerful paradigm for model alignment, yet its reliance on teacher logits restricts its application to white-box scenarios. We contend that struct…