2 papers
cs.LG2026
Distill-then-Replace: Efficient Task-Specific Hybrid Attention Model Construction
Xiaojie Xia, Huigang Zhang, Chaoliang Zhong +2
Transformer architectures deliver state-of-the-art accuracy via dense full-attention, but their quadratic time and memory complexity with respect to sequence length limits practica…
cs.AI2026
Following the Teacher's Footsteps: Scheduled Checkpoint Distillation for Domain-Specific LLMs
Cheng Feng, Chaoliang Zhong, Jun Sun +1
Large language models (LLMs) are challenging to deploy for domain-specific tasks due to their massive scale. While distilling a fine-tuned LLM into a smaller student model is a pro…