1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Sangkyu Lee, Janghoon Han, Hosung Song +3
Direct Preference Optimization (DPO) demonstrates the advantage of aligning a large language model with human preference using only an offline dataset. However, DPO has the limitat…