1 paper
Shuai Dong, Yongfu Zhu, Yuqi Xu +26
Reinforcement learning (RL) improves reasoning, but its performance depends on the policy from which training begins. We study on-policy distillation (OPD) as a preparation stage f…