3 papers
cs.CL2026
TALAS: Teacher-Anchored Layer Alignment with Adaptive Sharpness-Aware Minimization for Embedding Distillation
Quoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi +4
Knowledge Distillation (KD) has established itself as a pivotal technique for compressing large pre-trained language models. However, existing methods that force a student to stric…
cs.CL2026
MTA: Multi-Granular Trajectory Alignment for Large Language Model Distillation
Pham Khanh Chi, Quoc Phong Dao, Thuat Nguyen +3
Knowledge distillation is a key technique for compressing large language models (LLMs), but most existing methods align representations at fixed layers or token-level outputs, igno…
cs.CL2026
SRA: Span Representation Alignment for Large Language Model Distillation
Quoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi +4
Cross-Tokenizer Knowledge Distillation (CTKD) enables knowledge transfer between a large language model and a smaller student, even when they employ different tokenizers. While exi…