1 paper · 1 filter
Yuhao Li, Shengchao Liu
Debates about large language model post-training often treat supervised fine-tuning (SFT) as imitation and reinforcement learning (RL) as discovery. But this distinction is too coa…