4 papers
Beyond RLHF: A Unified Theoretical Framework of Alignment
Jihun Yun, Juno Kim, Jongho Park +4
Alignment via reinforcement learning from human feedback (RLHF) has become the dominant paradigm for controlling the quality of outputs from large language models (LLMs). However,…
Coverage Improvement and Fast Convergence of On-policy Preference Learning
Juno Kim, Jihun Yun, Jason D. Lee +1
Online on-policy preference learning algorithms for language model alignment such as online direct policy optimization (DPO) can significantly outperform their offline counterparts…
Improved Offline Contextual Bandits with Second-Order Bounds: Betting and Freezing
J. Jon Ryu, Jeongyeol Kwon, Benjamin Koppe +1
We consider off-policy selection and learning in contextual bandits, where the learner aims to select or train a reward-maximizing policy using data collected by a fixed behavior p…
Learning Explainable Dense Reward Shapes via Bayesian Optimization
Ryan Koo, Ian Yang, Vipul Raheja +3
Current reinforcement learning from human feedback (RLHF) pipelines for large language model (LLM) alignment typically assign scalar rewards to sequences, using the final token as…