1 paper · 1 filter
Huaye Zeng, Dongfu Jiang, Haozhe Wang +3
Most progress in recent coder models has been driven by supervised fine-tuning (SFT), while the potential of reinforcement learning (RL) remains largely unexplored, primarily due t…