2 papers
cs.CL2026
The Flip Side of RLHF: On-Policy Feedback for Reward Model Self-Supervised Improvement
Xiaobo Wang, Tong Wu, Min Tang +3
Building strong reward models (RMs) for language model alignment is bottlenecked by the cost and difficulty of acquiring diverse and reliable preference data from human annotation…
cs.LG2025
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models
Qi Liu, Jingqing Ruan, Hao Li +7
Existing multi-objective preference alignment methods for large language models (LLMs) face limitations: (1) the inability to effectively balance various preference dimensions, and…