18 citations · 18 across the 7 of their papers we have counts for
7 papers
LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL
Sijie Wang, Zhiqiang Tan, Xinrui Yang +1
Diffusion reinforcement learning (RL) has recently achieved significant success in post-training image and video generative models. However, most diffusion RL methods, including Da…
Unregularized Convergence of Single-Loop, Entropy-Regularized Natural Actor-Critic
Zhiqiang Tan
While entropy regularization is widely used to stabilize and accelerate Natural Policy Gradient methods, its ability to yield faster convergence rates for the unregularized objecti…
Bidirectional Resource Scheduling for Disaggregated and Asynchronous RL Post-Training
Zhiqiang Tan, Maoxin Wang, Sijie Wang +4
It is well established that the reasoning capabilities of large language models (LLMs) can be improved by applying reinforcement learning (RL) in a post-training stage. In a standa…
Accelerating Disaggregated RL for Visual Generative LLMs with Diffusion-Based Parallelism and Trainer-Assisted Generation
Sijie Wang, Zhengyu Qing, Zhiqiang Tan +6
Reinforcement learning (RL) has become a dominant post-training paradigm, driving the emergence of high-performance RL systems such as veRL for autoregressive large language models…
Preconditioned Discrete-HAMS: A Second-order Irreversible Discrete Sampler
Yuze Zhou, Zhiqiang Tan
Gradient-based Markov Chain Monte Carlo methods have recently received much attention for sampling discrete distributions, with notable examples such as Norm Constrained Gradient (…
Discrete Hamiltonian-Assisted Metropolis Sampling
Yuze Zhou, Zhiqiang Tan
Gradient-based Markov Chain Monte Carlo methods have recently received much attention for sampling discrete distributions, with interesting connections to their continuous counterp…