2 papers
cs.LG2024
VickreyFeedback: Cost-efficient Data Construction for Reinforcement Learning from Human Feedback
Guoxi Zhang, Jiuding Duan
This paper addresses the cost-efficiency aspect of Reinforcement Learning from Human Feedback (RLHF). RLHF leverages datasets of human preferences over outputs of large language mo…
cs.LG2024
A Generalized Model for Multidimensional Intransitivity
Jiuding Duan, Jiyi Li, Yukino Baba +1
Intransitivity is a critical issue in pairwise preference modeling. It refers to the intransitive pairwise preferences between a group of players or objects that potentially form a…