21 citations · 53 across the 6 of their papers we have counts for
8 papers
Multi-Task Off-Policy Learning from Bandit Feedback
Joey Hong, Branislav Kveton, Sumeet Katariya +2
Many practical applications, such as recommender systems and learning to rank, involve solving multiple similar tasks. One example is learning of recommendation policies for users…
When Should We Prefer Offline Reinforcement Learning Over Behavioral Cloning?
Aviral Kumar, Joey Hong, Anikait Singh +1
Offline reinforcement learning (RL) algorithms can acquire effective policies by utilizing previously collected experience, without any online interaction. It is widely understood…
Deep Hierarchy in Bandits
Joey Hong, Branislav Kveton, Sumeet Katariya +2
Mean rewards of actions are often correlated. The form of these correlations may be complex and unknown a priori, such as the preferences of a user for recommended products and the…
Non-Stationary Latent Bandits
Joey Hong, Branislav Kveton, Manzil Zaheer +4
Users of recommender systems often behave in a non-stationary fashion, due to their evolving preferences and tastes over time. In this work, we propose a practical approach for fas…
Latent Programmer: Discrete Latent Codes for Program Synthesis
Joey Hong, David Dohan, Rishabh Singh +2
In many sequence learning tasks, such as program synthesis and document summarization, a key problem is searching over a large space of possible output sequences. We propose to lea…
Latent Bandits Revisited
Joey Hong, Branislav Kveton, Manzil Zaheer +3
A latent bandit problem is one in which the learning agent knows the arm reward distributions conditioned on an unknown discrete latent state. The primary goal of the agent is to i…