Language Models are Few-Shot Learners
arXiv:2005.14165
Abstract
Recent work has demonstrated substantial gains on many NLP tasks and benchmarks by pre-training on a large corpus of text followed by fine-tuning on a specific task. While typically task-agnostic in architecture, this method still requires task-specific fine-tuning datasets of thousands or tens of thousands of examples. By contrast, humans can generally perform a new language task from only a few examples or from simple instructions - something which current NLP systems still largely struggle to do. Here we show that scaling up language models greatly improves task-agnostic, few-shot performance, sometimes even reaching competitiveness with prior state-of-the-art fine-tuning approaches. Specifically, we train GPT-3, an autoregressive language model with 175 billion parameters, 10x more than any previous non-sparse language model, and test its performance in the few-shot setting. For all tasks, GPT-3 is applied without any gradient updates or fine-tuning, with tasks and few-shot demonstrations specified purely via text interaction with the model. GPT-3 achieves strong performance on many NLP datasets, including translation, question-answering, and cloze tasks, as well as several tasks that require on-the-fly reasoning or domain adaptation, such as unscrambling words, using a novel word in a sentence, or performing 3-digit arithmetic. At the same time, we also identify some datasets where GPT-3's few-shot learning still struggles, as well as some datasets where GPT-3 faces methodological issues related to training on large web corpora. Finally, we find that GPT-3 can generate samples of news articles which human evaluators have difficulty distinguishing from articles written by humans. We discuss broader societal impacts of this finding and of GPT-3 in general.
40+32 pages
References in corpus (24)
- Distilling the Knowledge in a Neural Network
- NLTK: The Natural Language Toolkit
- Cross-lingual Language Model Pretraining
- Scaling Laws for Neural Language Models
- Multilingual Denoising Pre-training for Neural Machine Translation
- REALM: Retrieval-Augmented Language Model Pre-Training
- Generating Long Sequences with Sparse Transformers
- Deep Learning Scaling is Predictable, Empirically
- Multi-Task Deep Neural Networks for Natural Language Understanding
- BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
- Improving Multi-Task Deep Neural Networks via Knowledge Distillation for Natural Language Understanding
- Learning and Evaluating General Linguistic Intelligence
- Combining Independent Modules to Solve Multiple-choice Synonym and Analogy Problems
- DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs
- Adversarial Training for Large Neural Language Models
- Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference
- Generating Wikipedia by Summarizing Long Sequences
- Massively Multilingual Neural Machine Translation
- Technical report on Conversational Question Answering
- HellaSwag: Can a Machine Really Finish Your Sentence?
- SummAE: Zero-Shot Abstractive Text Summarization using Length-Agnostic Auto-Encoders
- Corpus-based Learning of Analogies and Semantic Relations
- Story Ending Prediction by Transferable BERT
- Reducing Gender Bias in Word-Level Language Models with a Gender-Equalizing Loss Function
Cited by in corpus (55)
- Linformer: Self-Attention with Linear Complexity
- GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding
- Carbontracker: Tracking and Predicting the Carbon Footprint of Training Deep Learning Models
- A Survey on Deep Learning for Localization and Mapping: Towards the Age of Spatial Machine Intelligence
- A Survey on Transfer Learning in Natural Language Processing
- Self-training Improves Pre-training for Natural Language Understanding
- Generative Language Modeling for Automated Theorem Proving
- learn2learn: A Library for Meta-Learning Research
- Language Models as Few-Shot Learner for Task-Oriented Dialogue Systems
- Controlling Style in Generated Dialogue
- Example-Based Named Entity Recognition
- Facts as Experts: Adaptable and Interpretable Neural Memory over Symbolic Knowledge
- Conditional Negative Sampling for Contrastive Learning of Visual Representations
- Neural Language Generation: Formulation, Methods, and Evaluation
- Auxiliary-task Based Deep Reinforcement Learning for Participant Selection Problem in Mobile Crowdsourcing
- A Survey on Self-supervised Pre-training for Sequential Transfer Learning in Neural Networks
- On Linear Identifiability of Learned Representations
- Robust Conversational AI with Grounded Text Generation
- PMI-Masking: Principled masking of correlated spans
- Transferring Inductive Biases through Knowledge Distillation
- Accelerating Sparse DNN Models without Hardware-Support via Tile-Wise Sparsity
- Mathematical Reasoning via Self-supervised Skip-tree Training
- Accenture at CheckThat! 2020: If you say so: Post-hoc fact-checking of claims using transformer-based models
- Periodic Stochastic Gradient Descent with Momentum for Decentralized Training
- Multi-node Bert-pretraining: Cost-efficient Approach
- Stochastic reserving with a stacked model based on a hybridized Artificial Neural Network
- MeLIME: Meaningful Local Explanation for Machine Learning Models
- PopMAG: Pop Music Accompaniment Generation
- Tearing Down the Memory Wall
- Surprisal-Triggered Conditional Computation with Neural Networks
- How To Evaluate Your Dialogue System: Probe Tasks as an Alternative for Token-level Evaluation Metrics
- Learning from Videos with Deep Convolutional LSTM Networks
- On Controllability of AI
- Train and You'll Miss It: Interactive Model Iteration with Weak Supervision and Pre-Trained Embeddings
- Tensorized Transformer for Dynamical Systems Modeling
- High-throughput relation extraction algorithm development associating knowledge articles and electronic health records
- Leam: An Interactive System for In-situ Visual Text Analysis
- Bespoke vs. Prêt-à-Porter Lottery Tickets: Exploiting Mask Similarity for Trainable Sub-Network Finding
- MLE-guided parameter search for task loss minimization in neural sequence modeling
- Knowledge Transfer via Pre-training for Recommendation: A Review and Prospect
- Weird AI Yankovic: Generating Parody Lyrics
- A Qualitative Evaluation of Language Models on Automatic Question-Answering for COVID-19
- Bollyrics: Automatic Lyrics Generator for Romanised Hindi
- Technical Report: Auxiliary Tuning and its Application to Conditional Text Generation
- Creative Captioning: An AI Grand Challenge Based on the Dixit Board Game
- A fast memoryless predictive algorithm in a chain of recurrent neural networks
- Navigating Human Language Models with Synthetic Agents
- The curious case of developmental BERTology: On sparsity, transfer learning, generalization and the brain
- Challenges and Thrills of Legal Arguments
- Computing Conceptual Distances between Breast Cancer Screening Guidelines: An Implementation of a Near-Peer Epistemic Model of Medical Disagreement
- Abstractive and mixed summarization for long-single documents
- End to End Dialogue Transformer
- Can questions summarize a corpus? Using question generation for characterizing COVID-19 research
- A Hybrid Natural Language Generation System Integrating Rules and Deep Learning Algorithms
- Hierarchical GPT with Congruent Transformers for Multi-Sentence Language Models