1 paper
Zehua Liu, Shuqi Liu, Tao Zhong +1
While Supervised Fine-Tuning (SFT) and Rejection Sampling Fine-Tuning (RFT) are standard for LLM alignment, they either rely on costly expert data or discard valuable negative samp…