111 citations · 158 across the 13 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Stabilizing Reinforcement Learning with LLMs: Formulation and Practices
Chujie Zheng, Kai Dang, Bowen Yu +10
This paper proposes a novel formulation for reinforcement learning (RL) with large language models, explaining why and under what conditions the true sequence-level reward can be o…
cs.LG2025
Soft Adaptive Policy Optimization
Chang Gao, Chujie Zheng, Xiong-Hui Chen +7
Reinforcement learning (RL) plays an increasingly important role in enhancing the reasoning capabilities of large language models (LLMs), yet stable and performant policy optimizat…
cs.LG2025
Group Sequence Policy Optimization
Chujie Zheng, Shixuan Liu, Mingze Li +9
This paper introduces Group Sequence Policy Optimization (GSPO), our stable, efficient, and performant reinforcement learning algorithm for training large language models. Unlike p…