2 papers
cs.AI2025
Preference Optimization via Contrastive Divergence: Your Reward Model is Secretly an NLL Estimator
Zhuotong Chen, Fang Liu, Xuan Zhu +2
Existing studies on preference optimization (PO) have centered on constructing pairwise preference data following simple heuristics, such as maximizing the margin between preferred…
cs.LG2024
Towards Improved Preference Optimization Pipeline: from Data Generation to Budget-Controlled Regularization
Zhuotong Chen, Fang Liu, Jennifer Zhu +2
Direct Preference Optimization (DPO) and its variants have become the de facto standards for aligning large language models (LLMs) with human preferences or specific goals. However…