K for the Price of 1: Parameter-efficient Multi-task and Transfer Learning
arXiv:1810.10703
Abstract
We introduce a novel method that enables parameter-efficient transfer and multi-task learning with deep neural networks. The basic approach is to learn a model patch - a small set of parameters - that will specialize to each task, instead of fine-tuning the last layer or the entire network. For instance, we show that learning a set of scales and biases is sufficient to convert a pretrained network to perform well on qualitatively different problems (e.g. converting a Single Shot MultiBox Detection (SSD) model into a 1000-class image classification model while reusing 98% of parameters of the SSD feature extractor). Similarly, we show that re-learning existing low-parameter layers (such as depth-wise convolutions) while keeping the rest of the network frozen also improves transfer-learning accuracy significantly. Our approach allows both simultaneous (multi-task) as well as sequential transfer learning. In several multi-task learning problems, despite using much fewer parameters than traditional logits-only fine-tuning, we match single-task performance.
published at ICLR 2019
References in corpus (12)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
- DeCAF: A Deep Convolutional Activation Feature for Generic Visual Recognition
- How transferable are features in deep neural networks?
- Federated Learning: Strategies for Improving Communication Efficiency
- Fine-Grained Visual Classification of Aircraft
- On the Number of Linear Regions of Deep Neural Networks
- Quantizing deep convolutional networks for efficient inference: A whitepaper
- Revisiting Batch Normalization For Practical Domain Adaptation
- Federated Meta-Learning with Fast Convergence and Efficient Communication
- ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design
Cited by in corpus (10)
- Tiny Machine Learning: Progress and Futures
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- Training BatchNorm and Only BatchNorm: On the Expressive Power of Random Features in CNNs
- TinyTL: Reduce Activations, Not Trainable Parameters for Efficient On-Device Learning
- ElasticTrainer: Speeding Up On-Device Training with Runtime Elastic Tensor Selection
- Multi-Task Federated Learning for Personalised Deep Neural Networks in Edge Computing
- Parameter-Efficient Transfer from Sequential Behaviors for User Modeling and Recommendation
- BasisNet: Two-stage Model Synthesis for Efficient Inference
- Resolution Switchable Networks for Runtime Efficient Image Recognition
- Learning Deep Multimodal Feature Representation with Asymmetric Multi-layer Fusion