49 citations · 101 across the 5 of their papers we have counts for
Showing stat.MLShow all
3 papers · 1 filter
stat.ML2016★ 49 cited
On Lower Bounds for Regret in Reinforcement Learning
Ian Osband, Benjamin Van Roy
This is a brief technical note to clarify the state of lower bounds on regret for reinforcement learning. In particular, this paper: - Reproduces a lower bound on regret for reinfo…
stat.ML2016★ 8 cited
Posterior Sampling for Reinforcement Learning Without Episodes
Ian Osband, Benjamin Van Roy
This is a brief technical note to clarify some of the issues with applying the application of the algorithm posterior sampling for reinforcement learning (PSRL) in environments wit…
stat.ML2014★ 43 cited
Model-based Reinforcement Learning and the Eluder Dimension
Ian Osband, Benjamin Van Roy
We consider the problem of learning to optimize an unknown Markov decision process (MDP). We show that, if the MDP can be parameterized within some known function class, we can obt…