1 paper
Tong Liu, Xiao Yu, Wenxuan Zhou +2
Efficient preference optimization algorithms such as Direct Preference Optimization (DPO) have become a popular approach in aligning large language models (LLMs) with human prefere…