1 paper · 1 filter
Ke Wang, Shuangqi Li, Mathieu Salzmann +1
Supervised fine-tuning (SFT) is an efficient approach for downstream task adaptation and often serves as the initialization stage for reinforcement learning (RL), but it can show w…