1 paper · 1 filter
Yifu Luo, Zeyu Chen, Haoyu Wang +4
On-policy self-distillation (OPSD) has proven effective for post-training large language models (LLMs), yet its application to diffusion LLMs (dLLMs) remains unexplored. Existing O…