Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
arXiv:2006.16236
Abstract
Transformers achieve remarkable performance in several tasks but due to their quadratic complexity, with respect to the input's length, they are prohibitively slow for very long sequences. To address this limitation, we express the self-attention as a linear dot-product of kernel feature maps and make use of the associativity property of matrix products to reduce the complexity from to , where is the sequence length. We show that this formulation permits an iterative implementation that dramatically accelerates autoregressive transformers and reveals their relationship to recurrent neural networks. Our linear transformers achieve similar performance to vanilla transformers and they are up to 4000x faster on autoregressive prediction of very long sequences.
ICML 2020, project at https://linear-transformers.com/
Cited by in corpus (38)
- CoAtNet: Marrying Convolution and Attention for All Data Sizes
- XCiT: Cross-Covariance Image Transformers
- Random Feature Attention
- Chasing Sparsity in Vision Transformers: An End-to-End Exploration
- Fastformer: Additive Attention Can Be All You Need
- LoFTR: Detector-Free Local Feature Matching with Transformers
- Luna: Linear Unified Nested Attention
- ATISS: Autoregressive Transformers for Indoor Scene Synthesis
- Container: Context Aggregation Network
- The Jazz Transformer on the Front Line: Exploring the Shortcomings of AI-composed Music through Quantitative Measures
- CCVS: Context-aware Controllable Video Synthesis
- Nyströmformer: A Nyström-Based Algorithm for Approximating Self-Attention
- Combiner: Full Attention Transformer with Sparse Computation Cost
- A Survey of Deep Learning Approaches for OCR and Document Understanding
- Centroid Transformers: Learning to Abstract with Attention
- Compound Word Transformer: Learning to Compose Full-Song Music over Dynamic Directed Hypergraphs
- STAR: Sparse Transformer-based Action Recognition
- Relative Positional Encoding for Transformers with Linear Complexity
- Going Beyond Linear Transformers with Recurrent Fast Weight Programmers
- Fast Convergence of DETR with Spatially Modulated Co-Attention
- Scatterbrain: Unifying Sparse and Low-rank Attention Approximation
- The Piano Inpainting Application
- Stable, Fast and Accurate: Kernelized Attention with Relative Positional Encoding
- Rethinking Lifelong Sequential Recommendation with Incremental Multi-Interest Attention
- FMMformer: Efficient and Flexible Transformer via Decomposed Near-field and Far-field Attention
- Not All Memories are Created Equal: Learning to Forget by Expiring
- TensorCoder: Dimension-Wise Attention via Tensor Representation for Natural Language Modeling
- Redesigning the Transformer Architecture with Insights from Multi-particle Dynamical Systems
- Transformed CNNs: recasting pre-trained convolutional layers with self-attention
- Smart Bird: Learnable Sparse Attention for Efficient and Effective Transformer
- Sparse Factorization of Large Square Matrices
- Beyond Nyströmformer -- Approximation of self-attention by Spectral Shifting
- Do Long-Range Language Models Actually Use Long-Range Context?
- Knowledge Transfer by Discriminative Pre-training for Academic Performance Prediction
- Poly-NL: Linear Complexity Non-local Layers with Polynomials
- Skyformer: Remodel Self-Attention with Gaussian Kernel and Nyström Method
- CoRe: An Efficient Coarse-refined Training Framework for BERT
- PermuteFormer: Efficient Relative Position Encoding for Long Sequences