Asymptotic theory of in-context learning by linear attention
arXiv:2405.11751 · doi:10.1073/pnas.2502599122
Abstract
Transformers have a remarkable ability to learn and execute tasks based on examples provided within the input itself, without explicit prior training. It has been argued that this capability, known as in-context learning (ICL), is a cornerstone of Transformers' success, yet questions about the necessary sample complexity, pretraining task diversity, and context length for successful ICL remain unresolved. Here, we provide a precise answer to these questions in an exactly solvable model of ICL of a linear regression task by linear attention. We derive sharp asymptotics for the learning curve in a phenomenologically-rich scaling regime where the token dimension is taken to infinity; the context length and pretraining task diversity scale proportionally with the token dimension; and the number of pretraining examples scales quadratically. We demonstrate a double-descent learning curve with increasing pretraining examples, and uncover a phase transition in the model's behavior between low and high task diversity regimes: In the low diversity regime, the model tends toward memorization of training tasks, whereas in the high diversity regime, it achieves genuine in-context learning and generalization beyond the scope of pretrained tasks. These theoretical insights are empirically validated through experiments with both linear attention and full nonlinear Transformer architectures.
15 pages (main doc), 6 figures, and supplementary information (22 pages)
References in corpus (26)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Emergent Abilities of Large Language Models
- Linformer: Self-Attention with Linear Complexity
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- Predictability and Surprise in Large Generative Models
- Are Emergent Abilities of Large Language Models a Mirage?
- Transformers learn in-context by gradient descent
- What learning algorithm is in-context learning? Investigations with linear models
- Learning curves of generic features maps for realistic datasets with a teacher-student model
- Data Distributional Properties Drive Emergent In-Context Learning in Transformers
- Scaling Data-Constrained Language Models
- Universality of empirical risk minimization
- Trained Transformers Learn Linear Models In-Context
- Hidden Progress in Deep Learning: SGD Learns Parities Near the Computational Limit
- Transformers as Algorithms: Generalization and Stability in In-context Learning
- Pretraining task diversity and the emergence of non-Bayesian in-context learning for regression
- Transformers as Statisticians: Provable In-Context Learning with In-Context Algorithm Selection
- Birth of a Transformer: A Memory Viewpoint
- Scaling and renormalization in high-dimensional regression
- Transformers learn to implement preconditioned gradient descent for in-context learning
- The Transient Nature of Emergent In-Context Learning in Transformers
- Nadaraya-Watson kernel smoothing as a random energy model
- Universality for the global spectrum of random inner-product kernel matrices in the polynomial regime
- The mechanistic basis of data dependence and abrupt learning in an in-context classification task
- Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear Regression
- How Many Pretraining Tasks Are Needed for In-Context Learning of Linear Regression?