Multitask Prompted Training Enables Zero-Shot Task Generalization
arXiv:2110.08207
Abstract
Large language models have recently been shown to attain reasonable zero-shot generalization on a diverse set of tasks (Brown et al., 2020). It has been hypothesized that this is a consequence of implicit multitask learning in language models' pretraining (Radford et al., 2019). Can zero-shot generalization instead be directly induced by explicit multitask learning? To test this question at scale, we develop a system for easily mapping any natural language tasks into a human-readable prompted form. We convert a large set of supervised datasets, each with multiple prompts with diverse wording. These prompted datasets allow for benchmarking the ability of a model to perform completely held-out tasks. We fine-tune a pretrained encoder-decoder model (Raffel et al., 2020; Lester et al., 2021) on this multitask mixture covering a wide variety of tasks. The model attains strong zero-shot performance on several standard datasets, often outperforming models up to 16x its size. Further, our approach attains strong performance on a subset of tasks from the BIG-bench benchmark, outperforming models up to 6x its size. All trained models are available at https://github.com/bigscience-workshop/t-zero and all prompts are available at https://github.com/bigscience-workshop/promptsource.
ICLR 2022 Spotlight (with extended discussion)
References in corpus (13)
- On the Opportunities and Risks of Foundation Models
- Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales
- SAMSum Corpus: A Human-annotated Dialogue Dataset for Abstractive Summarization
- BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
- True Few-Shot Learning with Language Models
- Quantifying the Carbon Emissions of Machine Learning
- Carbon Emissions and Large Neural Network Training
- Entailment as Few-Shot Learner
- Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference
- Calibrate Before Use: Improving Few-Shot Performance of Language Models
- Finetuned Language Models Are Zero-Shot Learners
- ANLIzing the Adversarial Natural Language Inference Dataset
- Datasets: A Community Library for Natural Language Processing
Cited by in corpus (32)
- When Large Language Models Meet Personalization: Perspectives of Challenges and Opportunities
- Large AI Models in Health Informatics: Applications, Challenges, and the Future
- Regression Transformer: Concurrent sequence regression and generation for molecular language modeling
- Large Language Models as Zero-Shot Conversational Recommenders
- Grammatical Error Correction: A Survey of the State of the Art
- Towards autonomous system: flexible modular production system enhanced with large language model agents
- Finetuned Language Models Are Zero-Shot Learners
- Learning from models beyond fine-tuning
- Generative Speech Recognition Error Correction with Large Language Models and Task-Activating Prompting
- ThoughtSource: A central hub for large language model reasoning data
- A Survey on Symbolic Knowledge Distillation of Large Language Models
- A Piece of Theatre: Investigating How Teachers Design LLM Chatbots to Assist Adolescent Cyberbullying Education
- Examining Inter-Consistency of Large Language Models Collaboration: An In-depth Analysis via Debate
- Do LLMs Understand Social Knowledge? Evaluating the Sociability of Large Language Models with SocKET Benchmark
- Multimodal Fusion and Vision-Language Models: A Survey for Robot Vision
- Large Language Models Can be Lazy Learners: Analyze Shortcuts in In-Context Learning
- Solving Math Word Problems via Cooperative Reasoning induced Language Models
- Match-Prompt: Improving Multi-task Generalization Ability for Neural Text Matching via Prompt Learning
- The Life Cycle of Knowledge in Big Language Models: A Survey
- BigBIO: A Framework for Data-Centric Biomedical Natural Language Processing
- Undesirable Biases in NLP: Addressing Challenges of Measurement
- Soft Prompt Decoding for Multilingual Dense Retrieval
- Mirror: A Natural Language Interface for Data Querying, Summarization, and Visualization
- CRUISE-Screening: Living Literature Reviews Toolbox
- Complex QA and language models hybrid architectures, Survey
- AutoConv: Automatically Generating Information-seeking Conversations with Large Language Models
- ChatGPT as a commenter to the news: can LLMs generate human-like opinions?
- Multi-Head Adapter Routing for Cross-Task Generalization
- TabGenie: A Toolkit for Table-to-Text Generation
- Are NLP Models Good at Tracing Thoughts: An Overview of Narrative Understanding
- Can open source large language models be used for tumor documentation in Germany? -- An evaluation on urological doctors' notes
- Model-Generated Pretraining Signals Improves Zero-Shot Generalization of Text-to-Text Transformers