Delay and Cooperation in Nonstochastic Bandits
arXiv:1602.04741
Abstract
We study networks of communicating learning agents that cooperate to solve a common nonstochastic bandit problem. Agents use an underlying communication network to get messages about actions selected by other agents, and drop messages that took more than hops to arrive, where is a delay parameter. We introduce \textsc{Exp3-Coop}, a cooperative version of the {\sc Exp3} algorithm and prove that with actions and agents the average per-agent regret after rounds is at most of order , where is the independence number of the -th power of the connected communication graph . We then show that for any connected graph, for the regret bound is , strictly better than the minimax regret for noncooperating agents. More informed choices of lead to bounds which are arbitrarily close to the full information minimax regret when is dense. When has sparse components, we show that a variant of \textsc{Exp3-Coop}, allowing agents to choose their parameters according to their centrality in , strictly improves the regret. Finally, as a by-product of our analysis, we provide the first characterization of the minimax regret for bandit learning with delay.
30 pages
References in corpus (6)
- Slow Learners are Fast
- Efficient learning by implicit exploration in bandit problems with side observations
- Efficient Optimal Learning for Contextual Bandits
- Distributed Delayed Stochastic Optimization
- Asynchronous stochastic convex optimization
- Explore no more: Improved high-probability regret bounds for non-stochastic bandits
Cited by in corpus (26)
- Secure Mobile Edge Computing in IoT via Collaborative Online Learning
- Game of Thrones: Fully Distributed Learning for Multi-Player Bandits
- Bandit Learning in Decentralized Matching Markets
- An Optimal Algorithm for Adversarial Bandits with Arbitrary Delays
- Multi-Player Bandits: The Adversarial Case
- Bandit Online Learning with Unknown Delays
- Cooperative Online Learning: Keeping your Neighbors Updated
- Competing Bandits in Matching Markets
- Kernel Methods for Cooperative Multi-Agent Contextual Bandits
- Collaborative Learning with Limited Interaction: Tight Bounds for Distributed Exploration in Multi-Armed Bandits
- Robust Multi-Agent Multi-Armed Bandits
- When to Call Your Neighbor? Strategic Communication in Cooperative Stochastic Bandits
- Randomized Allocation with Nonparametric Estimation for Contextual Multi-Armed Bandits with Delayed Rewards
- Collaborative Top Distribution Identifications with Limited Interaction
- Stochastic Bandits with Delayed Composite Anonymous Feedback
- Cooperation Speeds Surfing: Use Co-Bandit!
- Cache Replacement as a MAB with Delayed Feedback and Decaying Costs
- Banker Online Mirror Descent
- On No-Sensing Adversarial Multi-player Multi-armed Bandits with Collision Communications
- Adaptive Pricing in Insurance: Generalized Linear Models and Gaussian Process Regression Approaches
- Scale-Free Adversarial Multi-Armed Bandit with Arbitrary Feedback Delays
- PAC-Learning Uniform Ergodic Communicative Networks
- Nonstochastic Bandits and Experts with Arm-Dependent Delays
- Improving Social Welfare While Preserving Autonomy via a Pareto Mediator
- To update or not to update? Delayed Nonparametric Bandits with Randomized Allocation
- Unknown Delay for Adversarial Bandit Setting with Multiple Play