1 paper
Siyuan Gan, Yuhan Li, Xiran Wang +5
On-policy distillation (OPD) has recently emerged as a popular post-training paradigm for large language models (LLMs), providing an efficient way to transfer the knowledge and cap…