Piggyback: Adapting a Single Network to Multiple Tasks by Learning to Mask Weights
arXiv:1801.06519
Abstract
This work presents a method for adapting a single, fixed deep neural network to multiple tasks without affecting performance on already learned tasks. By building upon ideas from network quantization and pruning, we learn binary masks that piggyback on an existing network, or are applied to unmodified weights of that network to provide good performance on a new task. These masks are learned in an end-to-end differentiable fashion, and incur a low overhead of 1 bit per network parameter, per task. Even though the underlying network is fixed, the ability to mask individual weights allows for the learning of a large number of filters. We show performance comparable to dedicated fine-tuned networks for a variety of classification tasks, including those with large domain shifts from the initial task (ImageNet), and a variety of network architectures. Unlike prior work, we do not suffer from catastrophic forgetting or competition between tasks, and our performance is agnostic to task ordering. Code available at https://github.com/arunmallya/piggyback.
References in corpus (3)
Cited by in corpus (16)
- A Comprehensive Study of Class Incremental Learning Algorithms for Visual Tasks
- Learn to Grow: A Continual Structure Learning Framework for Overcoming Catastrophic Forgetting
- SpotTune: Transfer Learning through Adaptive Fine-tuning
- M2KD: Multi-model and Multi-level Knowledge Distillation for Incremental Learning
- Depthwise Convolution is All You Need for Learning Multiple Visual Domains
- Heterogeneous Multi-task Learning with Expert Diversity
- Self-Supervised Training Enhances Online Continual Learning
- End-to-End Multi-Task Learning with Attention
- Carousel Memory: Rethinking the Design of Episodic Memory for Continual Learning
- Feature Partitioning for Efficient Multi-Task Architectures
- Incremental Learning with Maximum Entropy Regularization: Rethinking Forgetting and Intransigence
- Continual Learning via Bit-Level Information Preserving
- Initial Classifier Weights Replay for Memoryless Class Incremental Learning
- Low-Complexity Probing via Finding Subnetworks
- Split-and-Bridge: Adaptable Class Incremental Learning within a Single Neural Network
- Coarse-To-Fine Incremental Few-Shot Learning