Efficient Continual Learning with Modular Networks and Task-Driven Priors
arXiv:2012.12631
Abstract
Existing literature in Continual Learning (CL) has focused on overcoming catastrophic forgetting, the inability of the learner to recall how to perform tasks observed in the past. There are however other desirable properties of a CL system, such as the ability to transfer knowledge from previous tasks and to scale memory and compute sub-linearly with the number of tasks. Since most current benchmarks focus only on forgetting using short streams of tasks, we first propose a new suite of benchmarks to probe CL algorithms across these new axes. Finally, we introduce a new modular architecture, whose modules represent atomic skills that can be composed to perform a certain task. Learning a task reduces to figuring out which past modules to re-use, and which new modules to instantiate to solve the current task. Our learning algorithm leverages a task-driven prior over the exponential search space of all possible ways to combine modules, enabling efficient learning on long streams of tasks. Our experiments show that this modular architecture and learning algorithm perform competitively on widely used CL benchmarks while yielding superior performance on the more challenging benchmarks we introduce in this work.
Accepted as a conference paper at ICLR 2021
References in corpus (6)
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- PathNet: Evolution Channels Gradient Descent in Super Neural Networks
- Efficient Lifelong Learning with A-GEM
- Three scenarios for continual learning
- Online Meta-Learning
- What and Where: Learn to Plug Adapters via NAS for Multi-Domain Learning
Cited by in corpus (7)
- Recent Advances of Continual Learning in Computer Vision: An Overview
- Class Gradient Projection For Continual Learning
- Modularity in Deep Learning: A Survey
- Continual Referring Expression Comprehension via Dual Modular Memorization
- CUCL: Codebook for Unsupervised Continual Learning
- GROWN: GRow Only When Necessary for Continual Learning
- Visually Grounded Continual Language Learning with Selective Specialization