1 paper · 1 filter
Han Xiao, Yifan Niu, Dongyi Liu +2
On-Policy Distillation (OPD) facilitates the transfer of knowledge from domain expert to student in the post-training phase of Large Language Models (LLMs). However, the supervisio…