2 papers
cs.CL2024
Self-Evolution Knowledge Distillation for LLM-based Machine Translation
Yuncheng Song, Liang Ding, Changtong Zan +1
Knowledge distillation (KD) has shown great promise in transferring knowledge from larger teacher models to smaller student models. However, existing KD strategies for large langua…
cs.LG2024
Densely Distilling Cumulative Knowledge for Continual Learning
Zenglin Shi, Pei Liu, Tong Su +4
Continual learning, involving sequential training on diverse tasks, often faces catastrophic forgetting. While knowledge distillation-based approaches exhibit notable success in pr…