Showing math.STShow all
3 papers · 1 filter
math.ST2020
Finite-Time Analysis of Round-Robin Kullback-Leibler Upper Confidence Bounds for Optimal Adaptive Allocation with Multiple Plays and Markovian Rewards
Vrettos Moulos
We study an extension of the classic stochastic multi-armed bandit problem which involves multiple plays and Markovian rewards in the rested bandits setting. In order to tackle thi…
math.ST2020
A Hoeffding Inequality for Finite State Markov Chains and its Applications to Markovian Bandits
Vrettos Moulos
This paper develops a Hoeffding inequality for the partial sums , where is an irreducible Markov chain on a finite state sp…
math.ST2019
Optimal Best Markovian Arm Identification with Fixed Confidence
Vrettos Moulos
We give a complete characterization of the sampling complexity of best Markovian arm identification in one-parameter Markovian bandit models. We derive instance specific nonasympto…