1 paper
Xiangchi Yuan, Dachuan Shi, Chunhui Zhang +4
Reinforcement learning (RL) is central to post-training, particularly for agentic models that require specialized reasoning behaviors. In this setting, model merging offers a pract…