1 paper
Han Xiao, Yifan Niu, Dongyi Liu +2
On-Policy Distillation (OPD) facilitates the transfer of knowledge from domain expert to student in the post-training phase of Large Language Models (LLMs). However, the supervisio…