15 citations · 18 across the 17 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Differential Information Distribution: A Bayesian Perspective on Direct Preference Optimization
Yunjae Won, Hyunji Lee, Hyeonbin Hwang +1
Direct Preference Optimization (DPO) has been widely used for aligning language models with human preferences in a supervised manner. However, several key questions remain unresolv…
cs.LG2024
Aligning Large Language Models by On-Policy Self-Judgment
Sangkyu Lee, Sungdong Kim, Ashkan Yousefpour +3
Existing approaches for aligning large language models with human preferences face a trade-off that requires a separate reward model (RM) for on-policy learning. In this paper, we…
cs.LG2024
Rethinking the Role of Proxy Rewards in Language Model Alignment
Sungdong Kim, Minjoon Seo
Learning from human feedback via proxy reward modeling has been studied to align Large Language Models (LLMs) with human values. However, achieving reliable training through that p…