Multiple-Play Bandits in the Position-Based Model
arXiv:1606.02448
Abstract
Sequentially learning to place items in multi-position displays or lists is a task that can be cast into the multiple-play semi-bandit setting. However, a major concern in this context is when the system cannot decide whether the user feedback for each item is actually exploitable. Indeed, much of the content may have been simply ignored by the user. The present work proposes to exploit available information regarding the display position bias under the so-called Position-based click model (PBM). We first discuss how this model differs from the Cascade model and its variants considered in several recent works on multiple-play bandits. We then provide a novel regret lower bound for this model as well as computationally efficient algorithms that display good empirical and theoretical performance.
References in corpus (5)
- Combinatorial Bandits Revisited
- Cascading Bandits: Learning to Rank in the Cascade Model
- Optimal Regret Analysis of Thompson Sampling in Stochastic Multi-armed Bandit Problem with Multiple Plays
- Lipschitz Bandits: Regret Lower Bounds and Optimal Algorithms
- DCM Bandits: Learning to Rank with Multiple Clicks
Cited by in corpus (6)
- Unifying Online and Counterfactual Learning to Rank
- Polynomial-time Algorithms for Multiple-arm Identification with Full-bandit Feedback
- Probabilistic Permutation Graph Search: Black-Box Optimization for Fairness in Ranking
- Beyond the Click-Through Rate: Web Link Selection with Multi-level Feedback
- Sparse Stochastic Bandits
- Contributions to Representation Learning with Graph Autoencoders and Applications to Music Recommendation