5 papers
Learning in Markovian bandits with non-observable states and constrained decision epochs
Thomas Hira, Victor Boone, Urtzi Ayesta +1
This paper studies the problem of regret minimization in Markovian bandits with \emph{non-observable states} and possibly \emph{constrained} decision epochs. The focus is restricte…
Logarithmic Regret of Exploration in Average Reward Markov Decision Processes
Victor Boone, Bruno Gaujal
In average reward Markov decision processes, state-of-the-art algorithms for regret minimization follow a well-established framework: They are model-based, optimistic and episodic.…
Towards Blackwell Optimality: Bellman Optimality Is All You Can Get
Victor Boone, Adrienne Tuynman
Although average gain optimality is a commonly adopted performance measure in Markov Decision Processes (MDPs), it is often too asymptotic. Further incorporating measures of immedi…
Asymptotically optimal regret in communicating Markov decision processes
Victor Boone
In this paper, we present a learning algorithm that achieves asymptotically optimal regret for Markov decision processes in average reward under a communicating assumption. That is…
The regret lower bound for communicating Markov Decision Processes
Victor Boone, Odalric-Ambrym Maillard
This paper is devoted to the extension of the regret lower bound beyond ergodic Markov decision processes (MDPs) in the problem dependent setting. While the regret lower bound for…