Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Preference Elicitation for Offline Reinforcement Learning
Alizée Pace, Bernhard Schölkopf, Gunnar Rätsch +1
Applying reinforcement learning (RL) to real-world problems is often made challenging by the inability to interact with the environment and the difficulty of designing reward funct…
cs.LG2024
Uncertainty-Penalized Direct Preference Optimization
Sam Houliston, Alizée Pace, Alexander Immer +1
Aligning Large Language Models (LLMs) to human preferences in content, style, and presentation is challenging, in part because preferences are varied, context-dependent, and someti…