Deep Neural Linear Bandits: Overcoming Catastrophic Forgetting through Likelihood Matching
arXiv:1901.08612
Abstract
We study the neural-linear bandit model for solving sequential decision-making problems with high dimensional side information. Neural-linear bandits leverage the representation power of deep neural networks and combine it with efficient exploration mechanisms, designed for linear contextual bandits, on top of the last hidden layer. Since the representation is being optimized during learning, information regarding exploration with "old" features is lost. Here, we propose the first limited memory neural-linear bandit that is resilient to this phenomenon, which we term catastrophic forgetting. We evaluate our method on a variety of real-world data sets, including regression, classification, and sentiment analysis, and observe that our algorithm is resilient to catastrophic forgetting and achieves superior performance.
References in corpus (5)
- Overcoming catastrophic forgetting in neural networks
- Learn What Not to Learn: Action Elimination with Deep Reinforcement Learning
- Shallow Updates for Deep Reinforcement Learning
- Visualizing Dynamics: from t-SNE to SEMI-MDPs
- Is a picture worth a thousand words? A Deep Multi-Modal Fusion Architecture for Product Classification in e-commerce
Cited by in corpus (8)
- Neural Thompson Sampling
- Neural Contextual Bandits with Deep Representation and Shallow Exploration
- Multi-facet Contextual Bandits: A Neural Network Perspective
- Regularized OFU: an Efficient UCB Estimator forNon-linear Contextual Bandit
- Deep Upper Confidence Bound Algorithm for Contextual Bandit Ranking of Information Selection
- Neural Contextual Bandits without Regret
- Online Forgetting Process for Linear Regression Models
- A Map of Bandits for E-commerce