6 papers
Refined PAC-Bayes Bounds for Offline Bandits
Amaury Gouverneur, Tobias J. Oechtering, Mikael Skoglund
In this paper, we present refined probabilistic bounds on empirical reward estimates for off-policy learning in bandit problems. We build on the PAC-Bayesian bounds from Seldin et…
An Information-Theoretic Analysis of Thompson Sampling with Infinite Action Spaces
Amaury Gouverneur, Borja Rodriguez Gálvez, Tobias Oechtering +1
This paper studies the Bayesian regret of the Thompson Sampling algorithm for bandit problems, building on the information-theoretic framework introduced by Russo and Van Roy (2015…
An Information-Theoretic Analysis of Thompson Sampling for Logistic Bandits
Amaury Gouverneur, Borja Rodríguez-Gálvez, Tobias J. Oechtering +1
We study the performance of the Thompson Sampling algorithm for logistic bandit problems. In this setting, an agent receives binary rewards with probabilities determined by a logis…
Information-Theoretic Minimax Regret Bounds for Reinforcement Learning based on Duality
Raghav Bongole, Amaury Gouverneur, Borja Rodríguez-Gálvez +2
We study agents acting in an unknown environment where the agent's goal is to find a robust policy. We consider robust policies as policies that achieve high cumulative rewards for…
Chained Information-Theoretic bounds and Tight Regret Rate for Linear Bandit Problems
Amaury Gouverneur, Borja Rodríguez-Gálvez, Tobias J. Oechtering +1
This paper studies the Bayesian regret of a variant of the Thompson-Sampling algorithm for bandit problems. It builds upon the information-theoretic framework of [Russo and Van Roy…
Optimal measurement budget allocation for particle filtering
Antoine Aspeel, Amaury Gouverneur, Raphaël M. Jungers +1
Particle filtering is a powerful tool for target tracking. When the budget for observations is restricted, it is necessary to reduce the measurements to a limited amount of samples…