Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Delta Knowledge Distillation for Large Language Models
Yihan Cao, Yanbin Kang, Zhengming Xing +1
Knowledge distillation (KD) is a widely adopted approach for compressing large neural networks by transferring knowledge from a large teacher model to a smaller student model. In t…
cs.CL2024
Instruction Mining: Instruction Data Selection for Tuning Large Language Models
Yihan Cao, Yanbin Kang, Chi Wang +1
Large language models (LLMs) are initially pretrained for broad capabilities and then finetuned with instruction-following datasets to improve their performance in interacting with…