Online learning with Corrupted context: Corrupted Contextual Bandits
arXiv:2006.15194
Abstract
We consider a novel variant of the contextual bandit problem (i.e., the multi-armed bandit with side-information, or context, available to a decision-maker) where the context used at each decision may be corrupted ("useless context"). This new problem is motivated by certain on-line settings including clinical trial and ad recommendation applications. In order to address the corrupted-context setting,we propose to combine the standard contextual bandit approach with a classical multi-armed bandit mechanism. Unlike standard contextual bandit methods, we are able to learn from all iteration, even those with corrupted context, by improving the computing of the expectation for each arm. Promising empirical results are obtained on several real-life datasets.
References in corpus (8)
- A Contextual-Bandit Approach to Personalized News Article Recommendation
- Analysis of Thompson Sampling for the multi-armed bandit problem
- Thompson Sampling for Contextual Bandits with Linear Payoffs
- On Upper-Confidence Bound Policies for Non-Stationary Bandit Problems
- A Survey on Practical Applications of Multi-Armed and Contextual Bandits
- Beyond Backprop: Online Alternating Minimization with Auxiliary Variables
- Interpretable Multi-Objective Reinforcement Learning through Policy Orchestration
- Unified Models of Human Behavioral Agents in Bandits, Contextual Bandits and RL
Cited by in corpus (6)
- Online Semi-Supervised Learning with Bandit Feedback
- Contextual Bandit with Missing Rewards
- Online learning in bandits with predicted context
- Spectral Clustering using Eigenspectrum Shape Based Nystrom Sampling
- Computing the Dirichlet-Multinomial Log-Likelihood Function
- Etat de l'art sur l'application des bandits multi-bras