1 citations · 1 across the 5 of their papers we have counts for
1 paper · 1 filter
Biqing Qi, Pengfei Li, Fangyuan Li +3
Direct Preference Optimization (DPO) improves the alignment of large language models (LLMs) with human values by training directly on human preference datasets, eliminating the nee…