1 paper
Yan Chen, Taojie Zhu, Meng Zhang +4
Continual supervised fine-tuning (SFT) is the de facto recipe for adapting large language models (LLMs) to a stream of downstream tasks, but it suffers from catastrophic forgetting…