Sequential Transfer in Multi-armed Bandit with Finite Set of Models
arXiv:1307.6887
Abstract
Learning from prior tasks and transferring that experience to improve future performance is critical for building lifelong learning agents. Although results in supervised and reinforcement learning show that transfer may significantly improve the learning performance, most of the literature on transfer is focused on batch learning tasks. In this paper we study the problem of \textit{sequential transfer in online learning}, notably in the multi-armed bandit framework, where the objective is to minimize the cumulative regret over a sequence of tasks by incrementally transferring knowledge from prior tasks. We introduce a novel bandit algorithm based on a method-of-moments approach for the estimation of the possible tasks and derive regret bounds for it.
References in corpus (3)
Cited by in corpus (16)
- Online Clustering of Bandits
- Sequential Transfer in Multi-armed Bandit with Finite Set of Models
- Safe Policy Search for Lifelong Reinforcement Learning with Sublinear Regret
- Bayesian decision-making under misspecified priors with applications to meta-learning
- Meta-Learning Bandit Policies by Gradient Ascent
- Meta-learning with Stochastic Linear Bandits
- Meta-Thompson Sampling
- Optimal Multitask Linear Regression and Contextual Bandits under Sparse Heterogeneity
- Hierarchical Bayesian Bandits
- An Asymptotically Optimal Primal-Dual Incremental Algorithm for Contextual Linear Bandits
- Self-Tuning Bandits over Unknown Covariate-Shifts
- Lifelong Learning in Multi-Armed Bandits
- A Novel Confidence-Based Algorithm for Structured Bandits
- No Regrets for Learning the Prior in Bandits
- Multitask Bandit Learning Through Heterogeneous Feedback Aggregation
- Recommending with Recommendations