2 papers
cs.DC2026
Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training
Tan Zhiqiang, Zhiqiang Tan, Maoxin Wang +11
It is well established that the reasoning capabilities of large language models (LLMs) can be improved by applying reinforcement learning (RL) in a post-training stage. In a standa…
cs.AI2026
Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation
Sijie Wang, Zhengyu Qing, Zhiqiang Tan +6
Reinforcement learning (RL) has become a dominant post-training paradigm, driving the emergence of high-performance RL systems such as veRL for autoregressive large language models…