9 citations · 16 across the 3 of their papers we have counts for
3 papers
cs.LG2024
RLHF and IIA: Perverse Incentives
Wanqiao Xu, Shi Dong, Xiuyuan Lu +3
Existing algorithms for reinforcement learning from human feedback (RLHF) can incentivize responses at odds with preferences because they are based on models that assume independen…
cs.LG2016★ 9 cited
Stochastic Rank-1 Bandits
Sumeet Katariya, Branislav Kveton, Csaba Szepesvari +2
We propose stochastic rank- bandits, a class of online learning problems where at each step a learning agent chooses a pair of row and column arms, and receives the product of t…
cs.LG2014★ 7 cited
Learning to Act Greedily: Polymatroid Semi-Bandits
Branislav Kveton, Zheng Wen, Azin Ashkan +1
Many important optimization problems, such as the minimum spanning tree and minimum-cost flow, can be solved optimally by a greedy method. In this work, we study a learning variant…