2 papers
cs.CL2026
Multi-Aspect Knowledge Distillation for Language Model with Low-rank Factorization
Zihe Liu, Yulong Mao, Jinan Xu +2
Knowledge distillation is an effective technique for pre-trained language model compression. However, existing methods only focus on the knowledge distribution among layers, which…
cs.CL2026
KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models
Songming Zhang, Xue Zhang, Tong Zhang +3
Knowledge distillation (KD) is an essential technique to compress large language models (LLMs) into smaller ones. However, despite the distinct roles of the student model and the t…