A survey on multi-player bandits
arXiv:2211.16275
Abstract
Due mostly to its application to cognitive radio networks, multiplayer bandits gained a lot of interest in the last decade. A considerable progress has been made on its theoretical aspect. However, the current algorithms are far from applicable and many obstacles remain between these theoretical results and a possible implementation of multiplayer bandits algorithms in real cognitive radio networks. This survey contextualizes and organizes the rich multiplayer bandits literature. In light of the existing works, some clear directions for future research appear. We believe that a further study of these different directions might lead to theoretical algorithms adapted to real-world situations.
final version, accepted at JMLR
References in corpus (7)
- Online Bandit Learning against an Adaptive Adversary: from Regret to Policy Regret
- Combinatorial semi-bandit with known covariance
- Multi-Player Bandits: The Adversarial Case
- Decentralized Multi-player Multi-armed Bandits with No Collision Information
- Decentralized Learning in Online Queuing Systems
- Asymptotically Optimal Strategies For Combinatorial Semi-Bandits in Polynomial Time
- Competing for Shareable Arms in Multi-Player Multi-Armed Bandits