Inductive Biases and Variable Creation in Self-Attention Mechanisms
arXiv:2110.10090
Abstract
Self-attention, an architectural motif designed to model long-range interactions in sequential data, has driven numerous recent breakthroughs in natural language processing and beyond. This work provides a theoretical analysis of the inductive biases of self-attention modules. Our focus is to rigorously establish which functions and long-range dependencies self-attention blocks prefer to represent. Our main result shows that bounded-norm Transformer networks "create sparse variables": a single self-attention head can represent a sparse function of the input sequence, with sample complexity scaling only logarithmically with the context length. To support our analysis, we present synthetic experiments to probe the sample complexity of learning sparse Boolean functions with Transformers.
v2: camera-ready revisions for ICML 2022
References in corpus (18)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Language Models are Few-Shot Learners
- Evaluating Large Language Models Trained on Code
- MLP-Mixer: An all-MLP Architecture for Vision
- Decision Transformer: Reinforcement Learning via Sequence Modeling
- Perceiver IO: A General Architecture for Structured Inputs & Outputs
- Spectrally-normalized margin bounds for neural networks
- Fantastic Generalization Measures and Where to Find Them
- Perceiver: General Perception with Iterative Attention
- Norm-Based Capacity Control in Neural Networks
- Generating Wikipedia by Summarizing Long Sequences
- Tensor Programs II: Neural Tangent Kernel for Any Architecture
- Generative Language Modeling for Automated Theorem Proving
- Infinite attention: NNGP and NTK for deep attention networks
- On Generalization Bounds of a Family of Recurrent Neural Networks
- A Mathematical Theory of Attention
- Statistically Meaningful Approximation: a Case Study on Approximating Turing Machines with Transformers
- Approximating How Single Head Attention Learns