1 paper · 1 filter
Keertana Chidambaram, Karthik Vinary Seetharaman, Vasilis Syrgkanis
Reinforcement Learning from Human Feedback (RLHF) has become central to aligning large language models with human values, typically by first learning a reward model from preference…