1 citations · 1 across the 2 of their papers we have counts for
1 paper · 1 filter
Ruohong Zhang, Liangke Gui, Zhiqing Sun +8
Preference modeling techniques, such as direct preference optimization (DPO), has shown effective in enhancing the generalization abilities of large language model (LLM). However,…