1 paper
Ke Zhang, Yunjie Tian, Dongdi Zhao +4
On-policy distillation (OPD), which supervises a student on its own sampled trajectories, has emerged as a data-efficient post-training method for improving reasoning while avoidin…