1 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Jie Liu, Zhanhui Zhou, Jiaheng Liu +4
Direct Preference Optimization (DPO), a standard method for aligning language models with human preferences, is traditionally applied to offline preferences. Recent studies show th…