1 paper
Hanbing Liu, Haoyang Li, Xiaokang Zhang +5
Direct Preference Optimization (DPO) has proven effective in complex reasoning tasks like math word problems and code generation. However, when applied to Text-to-SQL datasets, it…