3 citations · 3 across the 1 of their papers we have counts for
3 papers
cs.AI2024★ 3 cited
Direct Language Model Alignment from Online AI Feedback
Shangmin Guo, Biao Zhang, Tianlin Liu +9
Direct alignment from preferences (DAP) methods, such as DPO, have recently emerged as efficient alternatives to reinforcement learning from human feedback (RLHF), that do not requ…
cs.CL2024
LiPO: Listwise Preference Optimization through Learning-to-Rank
Tianqi Liu, Zhen Qin, Junru Wu +9
Aligning language models (LMs) with curated human feedback is critical to control their behaviors in real-world applications. Several recent policy optimization methods, such as DP…
cs.CL2023
Predicting Text Preference Via Structured Comparative Reasoning
Jing Nathan Yan, Tianqi Liu, Justin T Chiu +9
Comparative reasoning plays a crucial role in text preference prediction; however, large language models (LLMs) often demonstrate inconsistencies in their reasoning. While approach…