49 citations · 100 across the 3 of their papers we have counts for
5 papers
Approximate Thompson Sampling via Epistemic Neural Networks
Ian Osband, Zheng Wen, Seyed Mohammad Asghari +4
Thompson sampling (TS) is a popular heuristic for action selection, but it requires sampling from a posterior distribution. Unfortunately, this can become computationally intractab…
Robustness of Epinets against Distributional Shifts
Xiuyuan Lu, Ian Osband, Seyed Mohammad Asghari +4
Recent work introduced the epinet as a new approach to uncertainty modeling in deep learning. An epinet is a small neural network added to traditional neural networks, which, toget…
On Lower Bounds for Regret in Reinforcement Learning
Ian Osband, Benjamin Van Roy
This is a brief technical note to clarify the state of lower bounds on regret for reinforcement learning. In particular, this paper: - Reproduces a lower bound on regret for reinfo…
Posterior Sampling for Reinforcement Learning Without Episodes
Ian Osband, Benjamin Van Roy
This is a brief technical note to clarify some of the issues with applying the application of the algorithm posterior sampling for reinforcement learning (PSRL) in environments wit…
Model-based Reinforcement Learning and the Eluder Dimension
Ian Osband, Benjamin Van Roy
We consider the problem of learning to optimize an unknown Markov decision process (MDP). We show that, if the MDP can be parameterized within some known function class, we can obt…