Compacter: Efficient Low-Rank Hypercomplex Adapter Layers
arXiv:2106.04647
Abstract
Adapting large-scale pretrained language models to downstream tasks via fine-tuning is the standard method for achieving state-of-the-art performance on NLP benchmarks. However, fine-tuning all weights of models with millions or billions of parameters is sample-inefficient, unstable in low-resource settings, and wasteful as it requires storing a separate copy of the model for each task. Recent work has developed parameter-efficient fine-tuning methods, but these approaches either still require a relatively large number of parameters or underperform standard fine-tuning. In this work, we propose Compacter, a method for fine-tuning large-scale language models with a better trade-off between task performance and the number of trainable parameters than prior work. Compacter accomplishes this by building on top of ideas from adapters, low-rank optimization, and parameterized hypercomplex multiplication layers. Specifically, Compacter inserts task-specific weight matrices into a pretrained model's weights, which are computed efficiently as a sum of Kronecker products between shared "slow" weights and "fast" rank-one matrices defined per Compacter layer. By only training 0.047% of a pretrained model's parameters, Compacter performs on par with standard fine-tuning on GLUE and outperforms standard fine-tuning on SuperGLUE and low-resource settings. Our code is publicly available at~\url{https://github.com/rabeehk/compacter}.
accepted in NeurIPS, 2021
References in corpus (9)
- Linformer: Self-Attention with Linear Complexity
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- High-Performance Large-Scale Image Recognition Without Normalization
- Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
- BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
- True Few-Shot Learning with Language Models
- Parameter-Efficient Transfer Learning for NLP
- BatchEnsemble: An Alternative Approach to Efficient Ensemble and Lifelong Learning
- Beyond Fully-Connected Layers with Quaternions: Parameterization of Hypercomplex Multiplications with Parameters
Cited by in corpus (12)
- LoRA: Low-Rank Adaptation of Large Language Models
- Large Language Models Can Be Strong Differentially Private Learners
- Differentially Private Fine-tuning of Language Models
- Exploring Adapter-based Transfer Learning for Recommender Systems: Empirical Studies and Practical Insights
- IISAN: Efficiently Adapting Multimodal Representation for Sequential Recommendation with Decoupled PEFT
- PHNNs: Lightweight Neural Networks via Parameterized Hypercomplex Convolutions
- Visual Grounding with Multi-modal Conditional Adaptation
- Group-aware Parameter-efficient Updating for Content-Adaptive Neural Video Compression
- Category-wise Fine-Tuning: Resisting Incorrect Pseudo-Labels in Multi-Label Image Classification with Partial Labels
- Training Neural Networks with Fixed Sparse Masks
- The Efficiency Misnomer
- Efficient Attribute Injection for Pretrained Language Models