32 citations · 43 across the 4 of their papers we have counts for
7 papers · 1 filter
Autonomous exploration for navigating in non-stationary CMPs
Pratik Gajane, Ronald Ortner, Peter Auer +1
We consider a setting in which the objective is to learn to navigate in a controlled Markov process (CMP) where transition probabilities may abruptly change. For this setting, we p…
Variational Regret Bounds for Reinforcement Learning
Pratik Gajane, Ronald Ortner, Peter Auer
We consider undiscounted reinforcement learning in Markov decision processes (MDPs) where both the reward functions and the state-transition probabilities may vary (gradually or ab…
A Sliding-Window Algorithm for Markov Decision Processes with Arbitrarily Changing Rewards and Transitions
Pratik Gajane, Ronald Ortner, Peter Auer
We consider reinforcement learning in changing Markov Decision Processes where both the state-transition probabilities and the reward functions may vary over time. For this problem…
An algorithm with nearly optimal pseudo-regret for both stochastic and adversarial bandits
Peter Auer, Chao-Kai Chiang
We present an algorithm that achieves almost optimal pseudo-regret bounds against adversarial and stochastic bandits. Against adversarial bandits the pseudo-regret is $O(K\sqrt{n \…
Upper-Confidence-Bound Algorithms for Active Learning in Multi-Armed Bandits
Alexandra Carpentier, Alessandro Lazaric, Mohammad Ghavamzadeh +3
In this paper, we study the problem of estimating uniformly well the mean values of several distributions given a finite budget of samples. If the variance of the distributions wer…
Regret Bounds for Restless Markov Bandits
Ronald Ortner, Daniil Ryabko, Peter Auer +1
We consider the restless Markov bandit problem, in which the state of each arm evolves according to a Markov process independently of the learner's actions. We suggest an algorithm…