3 papers
cs.LG2025
Preference Elicitation for Offline Reinforcement Learning
Alizée Pace, Bernhard Schölkopf, Gunnar Rätsch +1
Applying reinforcement learning (RL) to real-world problems is often made challenging by the inability to interact with the environment and the difficulty of designing reward funct…
cs.LG2024
Uncertainty-Penalized Direct Preference Optimization
Sam Houliston, Alizée Pace, Alexander Immer +1
Aligning Large Language Models (LLMs) to human preferences in content, style, and presentation is challenging, in part because preferences are varied, context-dependent, and someti…
cs.CL2024
West-of-N: Synthetic Preferences for Self-Improving Reward Models
Alizée Pace, Jonathan Mallinson, Eric Malmi +2
The success of reinforcement learning from human feedback (RLHF) in language model alignment is strongly dependent on the quality of the underlying reward model. In this paper, we…