1 paper
Junming Yang, Ning Xu, Biao Liu +2
Preference optimization is crucial for aligning large language models (LLMs) with human values and intentions. A significant challenge in this process is the distribution mismatch…