1 paper · 1 filter
Xurong Xie, Zhucun Xue, Jiafu Wu +5
Knowledge distillation (KD) is a key technique for compressing large-scale language models (LLMs), yet prevailing logit-based methods typically employ static strategies that are mi…