10 citations · 12 across the 3 of their papers we have counts for
4 papers
Optimistic Posterior Sampling for Reinforcement Learning with Few Samples and Tight Guarantees
Daniil Tiapkin, Denis Belomestny, Daniele Calandriello +6
We consider reinforcement learning in an environment modeled by an episodic, finite, stage-dependent Markov decision process of horizon with states, and actions. The pe…
The Influence of Shape Constraints on the Thresholding Bandit Problem
James Cheshire, Pierre Menard, Alexandra Carpentier
We investigate the stochastic Thresholding Bandit problem (TBP) under several shape constraints. On top of (i) the vanilla, unstructured TBP, we consider the case where (ii) the se…
Non-Asymptotic Pure Exploration by Solving Games
Rémy Degenne, Wouter M. Koolen, Pierre Ménard
Pure exploration (aka active testing) is the fundamental task of sequentially gathering information to answer a query about a stochastic environment. Good algorithms make few mista…
Gradient Ascent for Active Exploration in Bandit Problems
Pierre Ménard
We present a new algorithm based on an gradient ascent for a general Active Exploration bandit problem in the fixed confidence setting. This problem encompasses several well studie…