5 papers · 1 filter
CPGD: Toward Stable Rule-based Reinforcement Learning for Language Models
Zongkai Liu, Fanqing Meng, Lingxiao Du +4
Recent advances in rule-based reinforcement learning (RL) have significantly improved the reasoning capability of language models (LMs) with rule-based rewards. However, existing R…
Rapid Learning in Constrained Minimax Games with Negative Momentum
Zijian Fang, Zongkai Liu, Chao Yu +1
In this paper, we delve into the utilization of the negative momentum technique in constrained minimax games. From an intuitive mechanical standpoint, we introduce a novel framewor…
An Offline Adaptation Framework for Constrained Multi-Objective Reinforcement Learning
Qian Lin, Zongkai Liu, Danying Mo +1
In recent years, significant progress has been made in multi-objective reinforcement learning (RL) research, which aims to balance multiple objectives by incorporating preferences…
Policy-regularized Offline Multi-objective Reinforcement Learning
Qian Lin, Chao Yu, Zongkai Liu +1
In this paper, we aim to utilize only offline trajectory data to train a policy for multi-objective RL. We extend the offline policy-regularized method, a widely-adopted approach f…
Off-Policy Primal-Dual Safe Reinforcement Learning
Zifan Wu, Bo Tang, Qian Lin +5
Primal-dual safe RL methods commonly perform iterations between the primal update of the policy and the dual update of the Lagrange Multiplier. Such a training paradigm is highly s…