1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.LG2024★ 1 cited
Offline Regularised Reinforcement Learning for Large Language Models Alignment
Pierre Harvey Richemond, Yunhao Tang, Daniel Guo +15
The dominant framework for alignment of large language models (LLM), whether through reinforcement learning from human feedback or direct preference optimisation, is to learn from…
cs.CL2024
West-of-N: Synthetic Preferences for Self-Improving Reward Models
Alizée Pace, Jonathan Mallinson, Eric Malmi +2
The success of reinforcement learning from human feedback (RLHF) in language model alignment is strongly dependent on the quality of the underlying reward model. In this paper, we…