activity
20112019
most citedRegret Bounds for Restless Markov Bandits

32 citations · 43 across the 4 of their papers we have counts for

collaborators
Showing cs.LGShow all

7 papers · 1 filter

cs.LG20197 cited

Autonomous exploration for navigating in non-stationary CMPs

Pratik Gajane, Ronald Ortner, Peter Auer +1

We consider a setting in which the objective is to learn to navigate in a controlled Markov process (CMP) where transition probabilities may abruptly change. For this setting, we p…

cs.LG2019

Variational Regret Bounds for Reinforcement Learning

Pratik Gajane, Ronald Ortner, Peter Auer

We consider undiscounted reinforcement learning in Markov decision processes (MDPs) where both the reward functions and the state-transition probabilities may vary (gradually or ab…

cs.LG2018

A Sliding-Window Algorithm for Markov Decision Processes with Arbitrarily Changing Rewards and Transitions

Pratik Gajane, Ronald Ortner, Peter Auer

We consider reinforcement learning in changing Markov Decision Processes where both the state-transition probabilities and the reward functions may vary over time. For this problem…

cs.LG2016

An algorithm with nearly optimal pseudo-regret for both stochastic and adversarial bandits

Peter Auer, Chao-Kai Chiang

We present an algorithm that achieves almost optimal pseudo-regret bounds against adversarial and stochastic bandits. Against adversarial bandits the pseudo-regret is $O(K\sqrt{n \…

cs.LG2015

Upper-Confidence-Bound Algorithms for Active Learning in Multi-Armed Bandits

Alexandra Carpentier, Alessandro Lazaric, Mohammad Ghavamzadeh +3

In this paper, we study the problem of estimating uniformly well the mean values of several distributions given a finite budget of samples. If the variance of the distributions wer…

cs.LG201232 cited

Regret Bounds for Restless Markov Bandits

Ronald Ortner, Daniil Ryabko, Peter Auer +1

We consider the restless Markov bandit problem, in which the state of each arm evolves according to a Markov process independently of the learner's actions. We suggest an algorithm…