2 papers
cs.CL2025
Uncovering the Impact of Chain-of-Thought Reasoning for Direct Preference Optimization: Lessons from Text-to-SQL
Hanbing Liu, Haoyang Li, Xiaokang Zhang +5
Direct Preference Optimization (DPO) has proven effective in complex reasoning tasks like math word problems and code generation. However, when applied to Text-to-SQL datasets, it…
cs.GT2025
Optimal Contest Design with Entry Restriction
Hanbing Liu, Ningyuan Li, Weian Li +2
This paper explores the design of contests involving contestants, focusing on how the designer decides on the number of contestants allowed and the prize structure with a fixed…