1 paper · 1 filter
Huimin Xu, Xin Mao, Feng-Lin Li +4
Direct Preference Optimization (DPO) often struggles with long-chain mathematical reasoning. Existing approaches, such as Step-DPO, typically improve this by focusing on the first…