6 citations · 13 across the 6 of their papers we have counts for
Showing 2022Show all
2 papers · 1 filter
cs.LG2022★ 1 cited
Multi-Task Off-Policy Learning from Bandit Feedback
Joey Hong, Branislav Kveton, Sumeet Katariya +2
Many practical applications, such as recommender systems and learning to rank, involve solving multiple similar tasks. One example is learning of recommendation policies for users…
cs.LG2022
Deep Hierarchy in Bandits
Joey Hong, Branislav Kveton, Sumeet Katariya +2
Mean rewards of actions are often correlated. The form of these correlations may be complex and unknown a priori, such as the preferences of a user for recommended products and the…