153 citations · 362 across the 29 of their papers we have counts for
7 papers · 1 filter
Single-Agent Policy Tree Search With Guarantees
Laurent Orseau, Levi H. S. Lelis, Tor Lattimore +1
We introduce two novel tree search algorithms that use a policy to guide search. The first algorithm is a best-first enumeration that uses a cost function that allows us to prove a…
Garbage In, Reward Out: Bootstrapping Exploration in Multi-Armed Bandits
Branislav Kveton, Csaba Szepesvari, Sharan Vaswani +3
We propose a bandit algorithm that explores by randomizing its history of rewards. Specifically, it pulls the arm with the highest mean reward in a non-parametric bootstrap sample…
Online Learning to Rank with Features
Shuai Li, Tor Lattimore, Csaba Szepesvári
We introduce a new model for online ranking in which the click probability factors into an examination and attractiveness function and the attractiveness function is a linear funct…
Linear Bandits with Stochastic Delayed Feedback
Claire Vernade, Alexandra Carpentier, Tor Lattimore +3
Stochastic linear bandits are a natural and well-studied model for structured exploration/exploitation problems and are widely used in applications such as online marketing and rec…
BubbleRank: Safe Online Learning to Re-Rank via Implicit Click Feedback
Chang Li, Branislav Kveton, Tor Lattimore +4
In this paper, we study the problem of safe online learning to re-rank, where user feedback is used to improve the quality of displayed lists. Learning to rank has traditionally be…
TopRank: A practical algorithm for online stochastic ranking
Tor Lattimore, Branislav Kveton, Shuai Li +1
Online learning to rank is a sequential decision-making problem where in each round the learning agent chooses a list of items and receives feedback in the form of clicks from the…