70 citations · 296 across the 39 of their papers we have counts for
1 paper · 2 filters
Sijie Wang, Zhengyu Qing, Zhiqiang Tan +6
Reinforcement learning (RL) has become a dominant post-training paradigm, driving the emergence of high-performance RL systems such as veRL for autoregressive large language models…