Sample Complexity of Multi-task Reinforcement Learning
arXiv:1309.6821
Abstract
Transferring knowledge across a sequence of reinforcement-learning tasks is challenging, and has a number of important applications. Though there is encouraging empirical evidence that transfer can improve performance in subsequent reinforcement-learning tasks, there has been very little theoretical analysis. In this paper, we introduce a new multi-task algorithm for a sequence of reinforcement-learning tasks when each task is sampled independently from (an unknown) distribution over a finite set of Markov decision processes whose parameters are initially unknown. For this setting, we prove under certain assumptions that the per-task sample complexity of exploration is reduced significantly due to transfer compared to standard single-task algorithms. Our multi-task algorithm also has the desired characteristic that it is guaranteed not to exhibit negative transfer: in the worst case its per-task sample complexity is comparable to the corresponding single-task algorithm.
Appears in Proceedings of the Twenty-Ninth Conference on Uncertainty in Artificial Intelligence (UAI2013)
References in corpus (1)
Cited by in corpus (23)
- Deep Decentralized Multi-task Multi-Agent Reinforcement Learning under Partial Observability
- Online Clustering of Bandits
- Sharing Knowledge in Multi-Task Deep Reinforcement Learning
- Bootstrap Latent-Predictive Representations for Multitask Reinforcement Learning
- Invariant Causal Prediction for Block MDPs
- Multi-task Deep Reinforcement Learning with PopArt
- Model-Free Non-Stationary RL: Near-Optimal Regret and Applications in Multi-Agent RL and Inventory Control
- Context-Aware Policy Reuse
- How Does an Approximate Model Help in Reinforcement Learning?
- When Is Generalizable Reinforcement Learning Tractable?
- Markov Decision Processes with Continuous Side Information
- When Simple Exploration is Sample Efficient: Identifying Sufficient Conditions for Random Exploration to Yield PAC RL Algorithms
- Model-Based Reinforcement Learning Exploiting State-Action Equivalence
- Collaborative Filtering Bandits
- Context-Aware Bandits
- TempLe: Learning Template of Transitions for Sample Efficient Multi-task RL
- Provably Efficient Multi-Task Reinforcement Learning with Model Transfer
- Learning Meta Representations for Agents in Multi-Agent Reinforcement Learning
- Multitask Bandit Learning Through Heterogeneous Feedback Aggregation
- Multitasking Inhibits Semantic Drift
- On mechanisms for transfer using landmark value functions in multi-task lifelong reinforcement learning
- Bayesian Policy Reuse
- Reinforcement Learning in Reward-Mixing MDPs