1 paper · 1 filter
Hao Wang, Hao Gu, Hongming Piao +6
The standard post-training recipe for large reasoning models, supervised fine-tuning followed by reinforcement learning (SFT-then-RL), may limit the benefits of the RL stage: while…