Showing stat.MLShow all
3 papers · 1 filter
stat.ML2025
An Information-Theoretic Analysis of Thompson Sampling for Logistic Bandits
Amaury Gouverneur, Borja RodrÃguez-Gálvez, Tobias J. Oechtering +1
We study the performance of the Thompson Sampling algorithm for logistic bandit problems. In this setting, an agent receives binary rewards with probabilities determined by a logis…
stat.ML2025
Refined PAC-Bayes Bounds for Offline Bandits
Amaury Gouverneur, Tobias J. Oechtering, Mikael Skoglund
In this paper, we present refined probabilistic bounds on empirical reward estimates for off-policy learning in bandit problems. We build on the PAC-Bayesian bounds from Seldin et…
stat.ML2025
An Information-Theoretic Analysis of Thompson Sampling with Infinite Action Spaces
Amaury Gouverneur, Borja Rodriguez Gálvez, Tobias Oechtering +1
This paper studies the Bayesian regret of the Thompson Sampling algorithm for bandit problems, building on the information-theoretic framework introduced by Russo and Van Roy (2015…