1 citations · 2 across the 6 of their papers we have counts for
1 paper · 1 filter
Xiao Cui, Mo Zhu, Yulei Qin +3
Knowledge distillation (KD) has become a prevalent technique for compressing large language models (LLMs). Existing KD methods are constrained by the need for identical tokenizers…