1 paper
Anhao Zhao, Haoran Xin, Yingqi Fan +3
Knowledge distillation is central to LLM post-training, yet its design space remains poorly understood, especially alongside reinforcement learning (RL). We show that the prevailin…