Attention is Not All You Need: Pure Attention Loses Rank Doubly Exponentially with Depth
arXiv:2103.03404
Abstract
Attention-based architectures have become ubiquitous in machine learning, yet our understanding of the reasons for their effectiveness remains limited. This work proposes a new way to understand self-attention networks: we show that their output can be decomposed into a sum of smaller terms, each involving the operation of a sequence of attention heads across layers. Using this decomposition, we prove that self-attention possesses a strong inductive bias towards "token uniformity". Specifically, without skip connections or multi-layer perceptrons (MLPs), the output converges doubly exponentially to a rank-1 matrix. On the other hand, skip connections and MLPs stop the output from degeneration. Our experiments verify the identified convergence phenomena on different variants of standard transformer architectures.
References in corpus (4)
Cited by in corpus (18)
- ViTGAN: Training GANs with Vision Transformers
- A Survey of Visual Transformers
- Refiner: Refining Self-attention for Vision Transformers
- SpectralFormer: Rethinking Hyperspectral Image Classification with Transformers
- MetaFormer Is Actually What You Need for Vision
- Analogous to Evolutionary Algorithm: Designing a Unified Sequence Model
- Not All Attention Is All You Need
- Can Vision Transformers Perform Convolution?
- On the Expressive Power of Self-Attention Matrices
- An Empirical Study: Extensive Deep Temporal Point Process
- Locally Enhanced Self-Attention: Combining Self-Attention and Convolution as Local and Context Terms
- GroupBERT: Enhanced Transformer Architecture with Efficient Grouped Structures
- Augmented Shortcuts for Vision Transformers
- Graph Conditioned Sparse-Attention for Improved Source Code Understanding
- Skyformer: Remodel Self-Attention with Gaussian Kernel and Nyström Method
- EfficientBERT: Progressively Searching Multilayer Perceptron via Warm-up Knowledge Distillation
- Multi-view 3D Reconstruction with Transformer
- Abstraction, Reasoning and Deep Learning: A Study of the "Look and Say" Sequence