Differentiable Prompt Makes Pre-trained Language Models Better Few-shot Learners
arXiv:2108.13161
Abstract
Large-scale pre-trained language models have contributed significantly to natural language processing by demonstrating remarkable abilities as few-shot learners. However, their effectiveness depends mainly on scaling the model parameters and prompt design, hindering their implementation in most real-world applications. This study proposes a novel pluggable, extensible, and efficient approach named DifferentiAble pRompT (DART), which can convert small language models into better few-shot learners without any prompt engineering. The main principle behind this approach involves reformulating potential natural language processing tasks into the task of a pre-trained language model and differentially optimizing the prompt template as well as the target label with backpropagation. Furthermore, the proposed approach can be: (i) Plugged to any pre-trained language models; (ii) Extended to widespread classification tasks. A comprehensive evaluation of standard NLP tasks demonstrates that the proposed approach achieves a better few-shot performance. Code is available in https://github.com/zjunlp/DART.
Accepted by ICLR 2022
References in corpus (11)
- KnowPrompt: Knowledge-aware Prompt-tuning with Synergistic Optimization for Relation Extraction
- Learning How to Ask: Querying LMs with Mixtures of Soft Prompts
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training
- Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
- True Few-Shot Learning with Language Models
- Making Pre-trained Language Models Better Few-shot Learners
- Entailment as Few-Shot Learner
- Calibrate Before Use: Improving Few-Shot Performance of Language Models
- PTR: Prompt Tuning with Rules for Text Classification
- Factual Probing Is [MASK]: Learning vs. Learning to Recall
Cited by in corpus (8)
- The Creativity of Text-to-Image Generation
- P-Tuning v2: Prompt Tuning Can Be Comparable to Fine-tuning Universally Across Scales and Tasks
- No More Fine-Tuning? An Experimental Evaluation of Prompt Tuning in Code Intelligence
- An Information-theoretic Approach to Prompt Engineering Without Ground Truth Labels
- The Rise of Artificial Intelligence in Educational Measurement: Opportunities and Ethical Challenges
- From Discrimination to Generation: Knowledge Graph Completion with Generative Transformer
- Ontology-enhanced Prompt-tuning for Few-shot Learning
- SentiPrompt: Sentiment Knowledge Enhanced Prompt-Tuning for Aspect-Based Sentiment Analysis