8 citations · 8 across the 2 of their papers we have counts for
3 papers
cs.LG2019
Accelerating the Computation of UCB and Related Indices for Reinforcement Learning
Wesley Cowan, Michael N. Katehakis, Daniel Pirutinsky
In this paper we derive an efficient method for computing the indices associated with an asymptotically optimal upper confidence bound algorithm (MDP-UCB) of Burnetas and Katehakis…
cs.LG2019
Reinforcement Learning: a Comparison of UCB Versus Alternative Adaptive Policies
Wesley Cowan, Michael N. Katehakis, Daniel Pirutinsky
In this paper we consider the basic version of Reinforcement Learning (RL) that involves computing optimal data driven (adaptive) policies for Markovian decision process with unkno…
stat.ML2015★ 8 cited
Normal Bandits of Unknown Means and Variances: Asymptotic Optimality, Finite Horizon Regret Bounds, and a Solution to an Open Problem
Wesley Cowan, Junya Honda, Michael N. Katehakis
Consider the problem of sampling sequentially from a finite number of populations, specified by random variables , and ; wh…