19 citations · 55 across the 11 of their papers we have counts for
5 papers · 1 filter
Sub-sampling for Efficient Non-Parametric Bandit Exploration
Dorian Baudry, Emilie Kaufmann, Odalric-Ambrym Maillard
In this paper we propose the first multi-armed bandit algorithm based on re-sampling that achieves asymptotically optimal regret simultaneously for different families of arms (name…
Episodic Reinforcement Learning in Finite MDPs: Minimax Lower Bounds Revisited
Omar Darwiche Domingues, Pierre Ménard, Emilie Kaufmann +1
In this paper, we propose new problem-independent lower bounds on the sample complexity and regret in episodic MDPs, with a particular focus on the non-stationary case in which the…
Fast active learning for pure exploration in reinforcement learning
Pierre Ménard, Omar Darwiche Domingues, Anders Jonsson +3
Realistic environments often provide agents with very limited feedback. When the environment is initially unknown, the feedback, in the beginning, can be completely absent, and the…
Planning in Markov Decision Processes with Gap-Dependent Sample Complexity
Anders Jonsson, Emilie Kaufmann, Pierre Ménard +3
We propose MDP-GapE, a new trajectory-based Monte-Carlo Tree Search algorithm for planning in a Markov Decision Process in which transitions have a finite support. We prove an uppe…
Adaptive Reward-Free Exploration
Emilie Kaufmann, Pierre Ménard, Omar Darwiche Domingues +3
Reward-free exploration is a reinforcement learning setting studied by Jin et al. (2020), who address it by running several algorithms with regret guarantees in parallel. In our wo…