1 paper · 1 filter
Duc Trung Vu, Pham Khanh Chi, Dat Phi Van +3
Knowledge Distillation (KD) has emerged as a crucial technique for compressing Large Language Models (LLMs). Although existing cross-tokenizer KD methods have made notable progress…