50 citations · 59 across the 2 of their papers we have counts for
6 papers
RecSim NG: Toward Principled Uncertainty Modeling for Recommender Ecosystems
Martin Mladenov, Chih-Wei Hsu, Vihan Jain +7
The development of recommender systems that optimize multi-turn interaction with users, and model the interactions of different agents (e.g., users, content providers, vendors) in…
Meta-Thompson Sampling
Branislav Kveton, Mikhail Konobeev, Manzil Zaheer +4
Efficient exploration in bandits is a fundamental online learning problem. We propose a variant of Thompson sampling that learns to explore better as it interacts with bandit insta…
Meta-Learning Bandit Policies by Gradient Ascent
Branislav Kveton, Martin Mladenov, Chih-Wei Hsu +3
Most bandit policies are designed to either minimize regret in any problem instance, making very few assumptions about the underlying environment, or in a Bayesian sense, assuming…
Differentiable Bandit Exploration
Craig Boutilier, Chih-Wei Hsu, Branislav Kveton +3
Exploration policies in Bayesian bandits maximize the average reward over problem instances drawn from some distribution . In this work, we learn such policies for an…
RecSim: A Configurable Simulation Platform for Recommender Systems
Eugene Ie, Chih-wei Hsu, Martin Mladenov +5
We propose RecSim, a configurable platform for authoring simulation environments for recommender systems (RSs) that naturally supports sequential interaction with users. RecSim all…
Empirical Bayes Regret Minimization
Chih-Wei Hsu, Branislav Kveton, Ofer Meshi +2
Most bandit algorithm designs are purely theoretical. Therefore, they have strong regret guarantees, but also are often too conservative in practice. In this work, we pioneer the i…