Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
Feng Luo, Yu-Neng Chuang, Guanchu Wang +4
On-policy distillation (OPD) trains student models under their own induced distribution while leveraging supervision from stronger teachers. We identify a failure mode of OPD: as t…
cs.CL2024
Dynamic Self-Distillation via Previous Mini-batches for Fine-tuning Small Language Models
Yao Fu, Yin Yu, Xiaotian Han +4
Knowledge distillation (KD) has become a widely adopted approach for compressing large language models (LLMs) to reduce computational costs and memory footprints. However, the avai…