Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024
Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization
Subhojyoti Mukherjee, Anusha Lalitha, Sailik Sengupta +2
Multi-objective alignment from human feedback (MOAHF) in large language models (LLMs) is a challenging problem as human preferences are complex, multifaceted, and often conflicting…
cs.LG2024
SeRA: Self-Reviewing and Alignment of Large Language Models using Implicit Reward Margins
Jongwoo Ko, Saket Dingliwal, Bhavana Ganesh +3
Direct alignment algorithms (DAAs), such as direct preference optimization (DPO), have become popular alternatives for Reinforcement Learning from Human Feedback (RLHF) due to thei…