5 citations · 6 across the 3 of their papers we have counts for
3 papers · 1 filter
SAIL: Self-Improving Efficient Online Alignment of Large Language Models
Mucong Ding, Souradip Chakraborty, Vibhu Agrawal +5
Reinforcement Learning from Human Feedback (RLHF) is a key method for aligning large language models (LLMs) with human preferences. However, current offline alignment approaches li…
PARL: A Unified Framework for Policy Alignment in Reinforcement Learning from Human Feedback
Souradip Chakraborty, Amrit Singh Bedi, Alec Koppel +4
We present a novel unified bilevel optimization-based framework, \textsf{PARL}, formulated to address the recently highlighted critical issue of policy alignment in reinforcement l…
On the Convergence and Sample Efficiency of Variance-Reduced Policy Gradient Method
Junyu Zhang, Chengzhuo Ni, Zheng Yu +2
Policy gradient (PG) gives rise to a rich class of reinforcement learning (RL) methods. Recently, there has been an emerging trend to accelerate the existing PG methods such as REI…