1 paper
Bowen Ding, Yuhan Chen, Jiayang Lyv +9
Supervised Fine-Tuning (SFT) and Reinforcement Learning (RL) dominate the post-training landscape for mathematical reasoning, yet differ fundamentally in their reliance on expert t…