4 papers
LC-GRPO: Bridging Train-Inference Gap for Flow-Based GRPO with Langevin Correction
Yingqing Guo, Hui Yuan, Zijian He +2
Flow-based generative models are typically sampled by solving a deterministic ordinary differential equation (ODE), whereas online reinforcement learning requires stochastic rollou…
TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion Models
Zheng Ding, Weirui Ye
Reinforcement learning (RL) post-training is crucial for aligning generative models with human preferences, but its prohibitive computational cost remains a major barrier to widesp…
C3Editor: Achieving Controllable Consistency in 2D Model for 3D Editing
Zeng Tao, Zheng Ding, Zeyuan Chen +3
Existing 2D-lifting-based 3D editing methods often encounter challenges related to inconsistency, stemming from the lack of view-consistent 2D editing models and the difficulty of…
Video2Policy: Scaling up Manipulation Tasks in Simulation through Internet Videos
Weirui Ye, Fangchen Liu, Zheng Ding +3
Simulation offers a promising approach for cheaply scaling training data for generalist policies. To scalably generate data from diverse and realistic tasks, existing algorithms ei…