1 citations · 1 across the 2 of their papers we have counts for
3 papers
cs.LG2025
On a Connection Between Imitation Learning and RLHF
Teng Xiao, Yige Yuan, Mingxiao Li +2
This work studies the alignment of large language models with preference data from an imitation learning perspective. We establish a close theoretical connection between reinforcem…
cs.LG2025
SimPER: A Minimalist Approach to Preference Alignment without Hyperparameters
Teng Xiao, Yige Yuan, Zhengyu Chen +4
Existing preference optimization objectives for language model alignment require additional hyperparameters that must be extensively tuned to achieve optimal performance, increasin…
cs.LG2024★ 1 cited
Cal-DPO: Calibrated Direct Preference Optimization for Language Model Alignment
Teng Xiao, Yige Yuan, Huaisheng Zhu +2
We study the problem of aligning large language models (LLMs) with human preference data. Contrastive preference optimization has shown promising results in aligning LLMs with avai…