3 papers
cs.AI2026
ShapE-GRPO: Shapley-Enhanced Reward Allocation for Multi-Candidate LLM Training
Rui Ai, Yu Pan, David Simchi-Levi +1
In user-agent interaction scenarios such as recommendation, brainstorming, and code suggestion, Large Language Models (LLMs) often generate sets of candidate recommendations where…
cs.LG2025
What Matters in Data for DPO?
Yu Pan, Zhongze Cai, Guanting Chen +2
Direct Preference Optimization (DPO) has emerged as a simple and effective approach for aligning large language models (LLMs) with human preferences, bypassing the need for a learn…
cs.LG2025
Understanding the Impact of Sampling Quality in Direct Preference Optimization
Kyung Rok Kim, Yumo Bai, Chonghuan Wang +1
We study how data of higher quality can be leveraged to improve performance in Direct Preference Optimization (DPO), aiming to understand its impact on DPO training dynamics. Our a…