1 citations · 2 across the 2 of their papers we have counts for
2 papers
cs.LG2023★ 1 cited
Confronting Reward Model Overoptimization with Constrained RLHF
Ted Moskovitz, Aaditya K. Singh, DJ Strouse +4
Large language models are typically aligned with human preferences by optimizing (RMs) fitted to human feedback. However, human preferences are multi-facet…
cs.LG2023★ 1 cited
Learning Optimal Advantage from Preferences and Mistaking it for Reward
W. Bradley Knox, Stephane Hatgis-Kessell, Sigurdur Orn Adalgeirsson +4
We consider algorithms for learning reward functions from human preferences over pairs of trajectory segments, as used in reinforcement learning from human feedback (RLHF). Most re…