178 citations · 427 across the 8 of their papers we have counts for
4 papers · 1 filter
On the Sample Complexity of Reinforcement Learning with a Generative Model
Mohammad Gheshlaghi Azar, Remi Munos, Bert Kappen
We consider the problem of learning the optimal action-value function in the discounted-reward Markov decision processes (MDPs). We prove a new PAC bound on the sample-complexity o…
Minimax Number of Strata for Online Stratified Sampling given Noisy Samples
Alexandra Carpentier, Rémi Munos
We consider the problem of online stratified sampling for Monte Carlo integration of a function given a finite budget of noisy evaluations to the function. More precisely we fo…
Bandit Theory meets Compressed Sensing for high dimensional Stochastic Linear Bandit
Alexandra Carpentier, Rémi Munos
We consider a linear stochastic bandit problem where the dimension of the unknown parameter is larger than the sampling budget . In such cases, it is in general impossib…
Thompson Sampling: An Asymptotically Optimal Finite Time Analysis
Emilie Kaufmann, Nathaniel Korda, Rémi Munos
The question of the optimality of Thompson Sampling for solving the stochastic multi-armed bandit problem had been open since 1933. In this paper we answer it positively for the ca…