5 papers · 1 filter
Learning What to Recommend: Minimax Optimal Simple Regret in Logistic Bandits
Shuai Liu, Alireza Bakhtiari, Alex Ayoub +2
We study stochastic logistic bandits with -dimensional action features under the simple-regret objective, where a learner uses rounds of exploration to output a single final…
Eluder dimension: localise it!
Alireza Bakhtiari, Alex Ayoub, Samuel Robertson +2
We establish a lower bound on the eluder dimension of generalised linear model classes, showing that standard eluder dimension-based analysis cannot lead to first-order regret boun…
Almost Free: Self-concordance in Natural Exponential Families and an Application to Bandits
Shuai Liu, Alex Ayoub, Flore Sentenac +2
We prove that single-parameter natural exponential families with subexponential tails are self-concordant with polynomial-sized parameters. For subgaussian natural exponential fami…
Switching the Loss Reduces the Cost in Batch (Offline) Reinforcement Learning
Alex Ayoub, Kaiwen Wang, Vincent Liu +5
We propose training fitted Q-iteration with log-loss (FQI-log) for batch reinforcement learning (RL). We show that the number of samples needed to learn a near-optimal policy with…
Exploration via linearly perturbed loss minimisation
David Janz, Shuai Liu, Alex Ayoub +1
We introduce exploration via linear loss perturbations (EVILL), a randomised exploration method for structured stochastic bandit problems that works by solving for the minimiser of…