1 paper · 1 filter
Weibin Liao, Xu Chu, Yasha Wang
In the domain of complex reasoning tasks, such as mathematical reasoning, recent advancements have proposed the use of Direct Preference Optimization (DPO) to suppress output of di…