1 paper
Haoru Tan, Sitong Wu, Yanfeng Chen +9
Reinforcement fine-tuning (RFT) is increasingly used to strengthen the reasoning abilities of large models, yet its effectiveness is bound by how training data are selected and use…