3 papers
cs.LG2023★ 2 cited
Behavior Alignment via Reward Function Optimization
Dhawal Gupta, Yash Chandak, Scott M. Jordan +2
Designing reward functions for efficiently guiding reinforcement learning (RL) agents toward specific behaviors is a complex task. This is challenging since it requires the identif…
cs.CL2023
Exploring the impact of low-rank adaptation on the performance, efficiency, and regularization of RLHF
Simeng Sun, Dhawal Gupta, Mohit Iyyer
During the last stage of RLHF, a large language model is aligned to human intents via PPO training, a process that generally requires large-scale computational resources. In this t…
cs.LG2023
Coagent Networks: Generalized and Scaled
James E. Kostas, Scott M. Jordan, Yash Chandak +5
Coagent networks for reinforcement learning (RL) [Thomas and Barto, 2011] provide a powerful and flexible framework for deriving principled learning rules for arbitrary stochastic…