6 citations · 6 across the 6 of their papers we have counts for
1 paper · 1 filter
Hanbing Liu, Haoyang Li, Xiaokang Zhang +5
Direct Preference Optimization (DPO) has proven effective in complex reasoning tasks like math word problems and code generation. However, when applied to Text-to-SQL datasets, it…