8 citations · 14 across the 21 of their papers we have counts for
3 papers · 2 filters
Scalable Representation Learning in Linear Contextual Bandits with Constant Regret Guarantees
Andrea Tirinzoni, Matteo Papini, Ahmed Touati +2
We study the problem of representation learning in stochastic contextual linear bandits. While the primary concern in this domain is usually to find realizable representations (i.e…
Online Learning with Off-Policy Feedback
Germano Gabbianelli, Matteo Papini, Gergely Neu
We study the problem of online learning in adversarial bandit problems under a partial observability model called off-policy feedback. In this sequential decision making problem, t…
Lifting the Information Ratio: An Information-Theoretic Analysis of Thompson Sampling for Contextual Bandits
Gergely Neu, Julia Olkhovskaya, Matteo Papini +1
We study the Bayesian regret of the renowned Thompson Sampling algorithm in contextual bandits with binary losses and adversarially-selected contexts. We adapt the information-theo…