A Survey on Contextual Multi-armed Bandits
arXiv:1508.03326
Abstract
In this survey we cover a few stochastic and adversarial contextual bandit algorithms. We analyze each algorithm's assumption and regret bound.
References in corpus (2)
Cited by in corpus (15)
- Online Learning: A Comprehensive Survey
- Algorithms with Logarithmic or Sublinear Regret for Constrained Contextual Bandits
- A Survey of Online Experiment Design with the Stochastic Multi-Armed Bandit
- Latent Contextual Bandits and their Application to Personalized Recommendations for New Users
- Metadata-based Multi-Task Bandits with Bayesian Hierarchical Models
- Contextual Constrained Learning for Dose-Finding Clinical Trials
- Decentralized Learning for Channel Allocation in IoT Networks over Unlicensed Bandwidth as a Contextual Multi-player Multi-armed Bandit Game
- Optimising Individual-Treatment-Effect Using Bandits
- Multi-Statistic Approximate Bayesian Computation with Multi-Armed Bandits
- A Multi-Armed Bandit-based Approach to Mobile Network Provider Selection
- Doubly Robust Interval Estimation for Optimal Policy Evaluation in Online Learning
- Learning Action Embeddings for Off-Policy Evaluation
- Learning to Decode: Reinforcement Learning for Decoding of Sparse Graph-Based Channel Codes
- Adversarial Linear Contextual Bandits with Graph-Structured Side Observations
- Rarely-switching linear bandits: optimization of causal effects for the real world