1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
Wenyi Xiao, Zechuan Wang, Leilei Gan +9
With the rapid advancement of large language models (LLMs), aligning policy models with human preferences has become increasingly critical. Direct Preference Optimization (DPO) has…