1 paper
Mingxu Tao, Chen Zhang, Quzhe Huang +4
Adapting large language models (LLMs) to new languages typically involves continual pre-training (CT) followed by supervised fine-tuning (SFT). However, this CT-then-SFT approach s…