1 citations · 1 across the 3 of their papers we have counts for
3 papers · 1 filter
A Decentralized Policy with Logarithmic Regret for a Class of Multi-Agent Multi-Armed Bandit Problems with Option Unavailability Constraints and Stochastic Communication Protocols
Pathmanathan Pankayaraj, D. H. S. Maithripala, J. M. Berg
This paper considers a multi-armed bandit (MAB) problem in which multiple mobile agents receive rewards by sampling from a collection of spatially dispersed stochastic processes, c…
A Decentralized Communication Policy for Multi Agent Multi Armed Bandit Problems
Pathmanathan Pankayaraj, D. H. S. Maithripala
This paper proposes a novel policy for a group of agents to, individually as well as collectively, solve a multi armed bandit (MAB) problem. The policy relies solely on the informa…
Asymptotic Allocation Rules for a Class of Dynamic Multi-armed Bandit Problems
T. W. U. Madhushani, D. H. S. Maithripala, N. E. Leonard
This paper presents a class of Dynamic Multi-Armed Bandit problems where the reward can be modeled as the noisy output of a time varying linear stochastic dynamic system that satis…