1 paper
Wei Chen, Yubing Wu, Junmei Yang +5
Preference optimization is widely used to align large language models (LLMs) with human preferences. However, many margin-based methods also suppress the chosen response when they…