2 papers
cs.LG2026
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
Mengqi Li, Lei Zhao, Anthony Man-Cho So +2
Can language models improve their reasoning performance without external rewards, using only their own sampled responses for training? We show that they can. We propose Self-evolvi…
cs.CL2025
DPO-Shift: Shifting the Distribution of Direct Preference Optimization
Xiliang Yang, Feng Jiang, Qianen Zhang +2
Direct Preference Optimization (DPO) and its variants have become increasingly popular for aligning language models with human preferences. These methods aim to teach models to bet…