2 papers
cs.CL2026
TALAS: Teacher-Anchored Layer Alignment with Adaptive Sharpness-Aware Minimization for Embedding Distillation
Quoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi +4
Knowledge Distillation (KD) has established itself as a pivotal technique for compressing large pre-trained language models. However, existing methods that force a student to stric…
cs.CL2026
SRA: Span Representation Alignment for Large Language Model Distillation
Quoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi +4
Cross-Tokenizer Knowledge Distillation (CTKD) enables knowledge transfer between a large language model and a smaller student, even when they employ different tokenizers. While exi…