2 papers
cs.CL2025
Delta Knowledge Distillation for Large Language Models
Yihan Cao, Yanbin Kang, Zhengming Xing +1
Knowledge distillation (KD) is a widely adopted approach for compressing large neural networks by transferring knowledge from a large teacher model to a smaller student model. In t…
cs.CL2025
A Survey on Post-training of Large Language Models
Guiyao Tie, Zeli Zhao, Dingjie Song +23
The emergence of Large Language Models (LLMs) has fundamentally transformed natural language processing, making them indispensable across domains ranging from conversational system…