2 citations · 7 across the 8 of their papers we have counts for
3 papers
A Provably Efficient Model-Free Posterior Sampling Method for Episodic Reinforcement Learning
Christoph Dann, Mehryar Mohri, Tong Zhang +1
Thompson Sampling is one of the most effective methods for contextual bandits and has been generalized to posterior sampling for certain MDP settings. However, existing posterior s…
Best of Both Worlds Model Selection
Aldo Pacchiano, Christoph Dann, Claudio Gentile
We study the problem of model selection in bandit scenarios in the presence of nested policy classes, with the goal of obtaining simultaneous adversarial and stochastic ("best of b…
Memory Lens: How Much Memory Does an Agent Use?
Christoph Dann, Katja Hofmann, Sebastian Nowozin
We propose a new method to study the internal memory used by reinforcement learning policies. We estimate the amount of relevant past information by estimating mutual information b…