1 paper
Haoyang Li, Sheng Lin, Fangcheng Fu +6
Reinforcement learning (RL) post-training has become pivotal for enhancing the capabilities of modern large models. A recent trend is to develop RL systems with a fully disaggregat…