1 paper · 1 filter
Moxin Li, Yuantao Zhang, Wenjie Wang +4
Multi-Objective Alignment (MOA) aims to align LLMs' responses with multiple human preference objectives, with Direct Preference Optimization (DPO) emerging as a prominent approach.…