1 citations · 1 across the 4 of their papers we have counts for
1 paper · 1 filter
Zekai Qu, Yinxu Pan, Ao Sun +2
Reinforcement learning (RL) post-training has become a trending paradigm for enhancing the capabilities of large language models (LLMs). Most existing RL systems for LLMs operate i…