Finetuned Language Models Are Zero-Shot Learners
arXiv:2109.01652
Abstract
This paper explores a simple method for improving the zero-shot learning abilities of language models. We show that instruction tuning -- finetuning language models on a collection of tasks described via instructions -- substantially improves zero-shot performance on unseen tasks. We take a 137B parameter pretrained language model and instruction-tune it on over 60 NLP tasks verbalized via natural language instruction templates. We evaluate this instruction-tuned model, which we call FLAN, on unseen task types. FLAN substantially improves the performance of its unmodified counterpart and surpasses zero-shot 175B GPT-3 on 20 of 25 tasks that we evaluate. FLAN even outperforms few-shot GPT-3 by a large margin on ANLI, RTE, BoolQ, AI2-ARC, OpenbookQA, and StoryCloze. Ablation studies reveal that number of finetuning datasets, model scale, and natural language instructions are key to the success of instruction tuning.
Version 5. Find list of changes in Appendix F (page 35)
References in corpus (21)
- Model-Agnostic Meta-Learning for Fast Adaptation of Deep Networks
- An Overview of Multi-Task Learning in Deep Neural Networks
- On the Opportunities and Risks of Foundation Models
- Evaluating Large Language Models Trained on Code
- Multilingual Denoising Pre-training for Neural Machine Translation
- Multitask Prompted Training Enables Zero-Shot Task Generalization
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Semi-supervised Sequence Learning
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Learning How to Ask: Querying LMs with Mixtures of Soft Prompts
- Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges
- ReCoRD: Bridging the Gap between Human and Machine Commonsense Reading Comprehension
- Meta-Learning: A Survey
- Multi-task learning for natural language processing in the 2020s: where are we going?
- Recursively Summarizing Books with Human Feedback
- Towards Zero-Label Language Learning
- Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets
- Program Synthesis with Large Language Models
- The Turking Test: Can Language Models Understand Instructions?
- Seq2Seq and Multi-Task Learning for joint intent and content extraction for domain specific interpreters
- CrossFit: A Few-shot Learning Challenge for Cross-task Generalization in NLP
Cited by in corpus (14)
- Multitask Prompted Training Enables Zero-Shot Task Generalization
- Recent Advances in Natural Language Processing via Large Pre-Trained Language Models: A Survey
- Large Language Models Can Be Strong Differentially Private Learners
- A General Language Assistant as a Laboratory for Alignment
- An Explanation of In-context Learning as Implicit Bayesian Inference
- NLP From Scratch Without Large-Scale Pretraining: A Simple and Efficient Framework
- Truthful AI: Developing and governing AI that does not lie
- ExT5: Towards Extreme Multi-Task Scaling for Transfer Learning
- Mirror: A Natural Language Interface for Data Querying, Summarization, and Visualization
- Curb Your Carbon Emissions: Benchmarking Carbon Emissions in Machine Translation
- Teaching Models new APIs: Domain-Agnostic Simulators for Task Oriented Dialogue
- CrossFit: A Few-shot Learning Challenge for Cross-task Generalization in NLP
- Cheap Learning: Maximising Performance of Language Models for Social Data Science Using Minimal Data
- A Survey on Green Deep Learning