13 citations · 13 across the 3 of their papers we have counts for
1 paper · 1 filter
Junkang Wu, Yuexiang Xie, Zhengyi Yang +6
This study addresses the challenge of noise in training datasets for Direct Preference Optimization (DPO), a method for aligning Large Language Models (LLMs) with human preferences…