2 papers
cs.LG2020
A Decentralized Policy with Logarithmic Regret for a Class of Multi-Agent Multi-Armed Bandit Problems with Option Unavailability Constraints and Stochastic Communication Protocols
Pathmanathan Pankayaraj, D. H. S. Maithripala, J. M. Berg
This paper considers a multi-armed bandit (MAB) problem in which multiple mobile agents receive rewards by sampling from a collection of spatially dispersed stochastic processes, c…
math.OC2019
Heterogeneous Stochastic Interactions for Multiple Agents in a Multi-armed Bandit Problem
Udari Madhushani, Naomi Ehrich Leonard
We define and analyze a multi-agent multi-armed bandit problem in which decision-making agents can observe the choices and rewards of their neighbors. Neighbors are defined by a ne…