Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
arXiv:2006.16236
Abstract
Transformers achieve remarkable performance in several tasks but due to their quadratic complexity, with respect to the input's length, they are prohibitively slow for very long sequences. To address this limitation, we express the self-attention as a linear dot-product of kernel feature maps and make use of the associativity property of matrix products to reduce the complexity from to , where is the sequence length. We show that this formulation permits an iterative implementation that dramatically accelerates autoregressive transformers and reveals their relationship to recurrent neural networks. Our linear transformers achieve similar performance to vanilla transformers and they are up to 4000x faster on autoregressive prediction of very long sequences.
ICML 2020, project at https://linear-transformers.com/
Cited by in corpus (14)
- Random Feature Attention
- LoFTR: Detector-Free Local Feature Matching with Transformers
- The Jazz Transformer on the Front Line: Exploring the Shortcomings of AI-composed Music through Quantitative Measures
- Nyströmformer: A Nyström-Based Algorithm for Approximating Self-Attention
- A Survey of Deep Learning Approaches for OCR and Document Understanding
- Centroid Transformers: Learning to Abstract with Attention
- Compound Word Transformer: Learning to Compose Full-Song Music over Dynamic Directed Hypergraphs
- Fast Convergence of DETR with Spatially Modulated Co-Attention
- Scatterbrain: Unifying Sparse and Low-rank Attention Approximation
- Redesigning the Transformer Architecture with Insights from Multi-particle Dynamical Systems
- TensorCoder: Dimension-Wise Attention via Tensor Representation for Natural Language Modeling
- Beyond Nyströmformer -- Approximation of self-attention by Spectral Shifting
- Skyformer: Remodel Self-Attention with Gaussian Kernel and Nyström Method
- CoRe: An Efficient Coarse-refined Training Framework for BERT