1 paper · 1 filter
He Zhu, Junyou Su, Peng Lai +4
Post-training of large language models involves a fundamental trade-off between supervised fine-tuning (SFT), which efficiently mimics demonstrations but tends to memorize, and rei…