Why Do Pretrained Language Models Help in Downstream Tasks? An Analysis of Head and Prompt Tuning
arXiv:2106.09226
Abstract
Pretrained language models have achieved state-of-the-art performance when adapted to a downstream NLP task. However, theoretical analysis of these models is scarce and challenging since the pretraining and downstream tasks can be very different. We propose an analysis framework that links the pretraining and downstream tasks with an underlying latent variable generative model of text -- the downstream classifier must recover a function of the posterior distribution over the latent variables. We analyze head tuning (learning a classifier on top of the frozen pretrained model) and prompt tuning in this setting. The generative model in our analysis is either a Hidden Markov Model (HMM) or an HMM augmented with a latent memory component, motivated by long-term dependencies in natural language. We show that 1) under certain non-degeneracy conditions on the HMM, simple classification heads can solve the downstream task, 2) prompt tuning obtains downstream guarantees with weaker non-degeneracy conditions, and 3) our recovery guarantees for the memory-augmented HMM are stronger than for the vanilla HMM because task-relevant information is easier to recover from the long-term memory. Experiments on synthetically generated data from HMMs back our theoretical findings.
References in corpus (5)
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- Theoretical Analysis of Self-Training with Deep Networks on Unlabeled Data
- Are Pre-trained Language Models Aware of Phrases? Simple but Strong Baselines for Grammar Induction
- Contrastive estimation reveals topic posterior information to linear models
- PMI-Masking: Principled masking of correlated spans
Cited by in corpus (6)
- On the Opportunities and Risks of Foundation Models
- Self-supervised Learning is More Robust to Dataset Imbalance
- An Explanation of In-context Learning as Implicit Bayesian Inference
- The Dark Side of the Language: Pre-trained Transformers in the DarkNet
- On a Benefit of Mask Language Modeling: Robustness to Simplicity Bias
- Analysing the Effect of Masking Length Distribution of MLM: An Evaluation Framework and Case Study on Chinese MRC Datasets