Multi-Player Bandits Revisited
arXiv:1711.02317
Abstract
Multi-player Multi-Armed Bandits (MAB) have been extensively studied in the literature, motivated by applications to Cognitive Radio systems. Driven by such applications as well, we motivate the introduction of several levels of feedback for multi-player MAB algorithms. Most existing work assume that sensing information is available to the algorithm. Under this assumption, we improve the state-of-the-art lower bound for the regret of any decentralized algorithms and introduce two algorithms, RandTopM and MCTopM, that are shown to empirically outperform existing algorithms. Moreover, we provide strong theoretical guarantees for these algorithms, including a notion of asymptotic optimality in terms of the number of selections of bad arms. We then introduce a promising heuristic, called Selfish, that can operate without sensing information, which is crucial for emerging applications to Internet of Things networks. We investigate the empirical performance of this algorithm and provide some first theoretical elements for the understanding of its behavior.
Cited by in corpus (13)
- A Practical Algorithm for Multiplayer Bandits when Arm Means Vary Among Players
- SIC-MMAB: Synchronisation Involves Communication in Multiplayer Multi-Armed Bandits
- Federated Linear Contextual Bandits
- Decentralized Multi-player Multi-armed Bandits with No Collision Information
- Selfish Robustness and Equilibria in Multi-Player Bandits
- Dominate or Delete: Decentralized Competing Bandits in Serial Dictatorship
- Online Learning for Cooperative Multi-Player Multi-Armed Bandits
- Distributed learning in congested environments with partial information
- My Fair Bandit: Distributed Learning of Max-Min Fairness with Multi-player Bandits
- Observe Before Play: Multi-armed Bandit with Pre-observations
- On No-Sensing Adversarial Multi-player Multi-armed Bandits with Collision Communications
- A High Performance, Low Complexity Algorithm for Multi-Player Bandits Without Collision Sensing Information
- GNU Radio Implementation of MALIN: "Multi-Armed bandits Learning for Internet-of-things Networks"