3 papers
stat.ML2024
Chained Information-Theoretic bounds and Tight Regret Rate for Linear Bandit Problems
Amaury Gouverneur, Borja Rodríguez-Gálvez, Tobias J. Oechtering +1
This paper studies the Bayesian regret of a variant of the Thompson-Sampling algorithm for bandit problems. It builds upon the information-theoretic framework of [Russo and Van Roy…
stat.ML2023
Thompson Sampling Regret Bounds for Contextual Bandits with sub-Gaussian rewards
Amaury Gouverneur, Borja Rodríguez-Gálvez, Tobias J. Oechtering +1
In this work, we study the performance of the Thompson Sampling algorithm for Contextual Bandit problems based on the framework introduced by Neu et al. and their concept of lifted…
cs.LG2022
An Information-Theoretic Analysis of Bayesian Reinforcement Learning
Amaury Gouverneur, Borja Rodríguez-Gálvez, Tobias J. Oechtering +1
Building on the framework introduced by Xu and Raginksy [1] for supervised learning problems, we study the best achievable performance for model-based Bayesian reinforcement learni…