Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
Reward-free Alignment for Conflicting Objectives
Peter Chen, Xiaopeng Li, Xi Chen +1
Direct alignment methods are increasingly used to align large language models (LLMs) with human preferences. However, many real-world alignment problems involve multiple conflictin…
cs.CL2025
ComPO: Preference Alignment via Comparison Oracles
Peter Chen, Xi Chen, Wotao Yin +1
Direct alignment methods are increasingly used for aligning large language models (LLMs) with human preferences. However, these methods suffer from the issues of verbosity and like…