A Survey on Practical Applications of Multi-Armed and Contextual Bandits
arXiv:1904.10040
Abstract
In recent years, multi-armed bandit (MAB) framework has attracted a lot of attention in various applications, from recommender systems and information retrieval to healthcare and finance, due to its stellar performance combined with certain attractive properties, such as learning from less feedback. The multi-armed bandit field is currently flourishing, as novel problem settings and algorithms motivated by various practical applications are being introduced, building on top of the classical bandit problem. This article aims to provide a comprehensive review of top recent developments in multiple real-life applications of the multi-armed bandit. Specifically, we introduce a taxonomy of common MAB-based applications and summarize state-of-art for each of those domains. Furthermore, we identify important current trends and provide new perspectives pertaining to the future of this exciting and fast-growing field.
under review by IJCAI 2019 Survey
References in corpus (1)
Cited by in corpus (21)
- Apollo: Transferable Architecture Exploration
- Fast and Interpretable Consensus Clustering via Minipatch Learning
- Regime Switching Bandits
- Online learning with Corrupted context: Corrupted Contextual Bandits
- Leveraging Good Representations in Linear Contextual Bandits
- Contextual Bandit with Missing Rewards
- An Asymptotically Optimal Primal-Dual Incremental Algorithm for Contextual Linear Bandits
- Feature Selection for Huge Data via Minipatch Learning
- Lifelong Learning in Multi-Armed Bandits
- Show Me the Whole World: Towards Entire Item Space Exploration for Interactive Personalized Recommendations
- Incentivized Bandit Learning with Self-Reinforcing User Preferences
- Solving Multi-Arm Bandit Using a Few Bits of Communication
- Computing the Dirichlet-Multinomial Log-Likelihood Function
- New Classes of the Greedy-Applicable Arm Feature Distributions in the Sparse Linear Bandit Problem
- Thompson Sampling via Local Uncertainty
- Etat de l'art sur l'application des bandits multi-bras
- Towards Intelligent Reconfigurable Wireless Physical Layer (PHY)
- Risk averse non-stationary multi-armed bandits
- A Map of Bandits for E-commerce
- Recurrent Neural-Linear Posterior Sampling for Nonstationary Contextual Bandits
- Intelligent and Reconfigurable Architecture for KL Divergence Based Online Machine Learning Algorithm