1 paper · 1 filter
Runming Yang, Taiqiang Wu, Jiahao Wang +4
Knowledge distillation (KD) has been a predominant method for compressing Large Language Models (LLMs). In this paper, we first revisit KD and Low-Rank Adaption (LoRA) and demonstr…