Exploring Versatile Generative Language Model Via Parameter-Efficient Transfer Learning
arXiv:2004.03829
Abstract
Fine-tuning pre-trained generative language models to down-stream language generation tasks has shown promising results. However, this comes with the cost of having a single, large model for each task, which is not ideal in low-memory/power scenarios (e.g., mobile). In this paper, we propose an effective way to fine-tune multiple down-stream generation tasks simultaneously using a single, large pre-trained model. The experiments on five diverse language generation tasks show that by just using an additional 2-3% parameters for each task, our model can maintain or even improve the performance of fine-tuning the whole model.
Accepted as Findings of EMNLP 2020, Zhaojiang Lin and Andrea Madotto contributed equally to this work
References in corpus (11)
- Distilling the Knowledge in a Neural Network
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Theoretical Models of Learning to Learn
- Pay Less Attention with Lightweight and Dynamic Convolutions
- TransferTransfo: A Transfer Learning Approach for Neural Network Based Conversational Agents
- An Actor-Critic Algorithm for Sequence Prediction
- Evaluating the State-of-the-Art of End-to-End Natural Language Generation: The E2E NLG Challenge
- Plug and Play Language Models: A Simple Approach to Controlled Text Generation
- BERT and PALs: Projected Attention Layers for Efficient Adaptation in Multi-Task Learning
- Understanding Knowledge Distillation in Non-autoregressive Machine Translation
- Shaping the Narrative Arc: An Information-Theoretic Approach to Collaborative Dialogue