52 citations · 52 across the 5 of their papers we have counts for
1 paper · 1 filter
Yuqian Fu, Tinghong Chen, Jiajun Chai +7
Large language models (LLMs) have achieved remarkable progress in reasoning tasks, yet the optimal integration of Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) remai…