Linformer: Self-Attention with Linear Complexity
arXiv:2006.04768
Abstract
Large transformer models have shown extraordinary success in achieving state-of-the-art results in many natural language processing applications. However, training and deploying these models can be prohibitively costly for long sequences, as the standard self-attention mechanism of the Transformer uses time and space with respect to sequence length. In this paper, we demonstrate that the self-attention mechanism can be approximated by a low-rank matrix. We further exploit this finding to propose a new self-attention mechanism, which reduces the overall self-attention complexity from to in both time and space. The resulting linear transformer, the \textit{Linformer}, performs on par with standard Transformer models, while being much more memory- and time-efficient.
References in corpus (5)
Cited by in corpus (73)
- A General Survey on Attention Mechanisms in Deep Learning
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting
- XCiT: Cross-Covariance Image Transformers
- Long Range Arena: A Benchmark for Efficient Transformers
- Multi-Behavior Hypergraph-Enhanced Transformer for Sequential Recommendation
- A Survey on Aspect-Based Sentiment Classification
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- CenterCLIP: Token Clustering for Efficient Text-Video Retrieval
- Random Feature Attention
- Chasing Sparsity in Vision Transformers: An End-to-End Exploration
- Fastformer: Additive Attention Can Be All You Need
- IA-RED: Interpretability-Aware Redundancy Reduction for Vision Transformers
- Multiscale Vision Transformers
- UTNet: A Hybrid Transformer Architecture for Medical Image Segmentation
- Luna: Linear Unified Nested Attention
- Keeping Your Eye on the Ball: Trajectory Attention in Video Transformers
- Container: Context Aggregation Network
- HTLM: Hyper-Text Pre-Training and Prompting of Language Models
- Nyströmformer: A Nyström-Based Algorithm for Approximating Self-Attention
- Combiner: Full Attention Transformer with Sparse Computation Cost
- Centroid Transformers: Learning to Abstract with Attention
- Conformer-Kernel with Query Term Independence for Document Retrieval
- Pre-Trained Models: Past, Present and Future
- Relative Positional Encoding for Transformers with Linear Complexity
- Fast Convergence of DETR with Spatially Modulated Co-Attention
- Vision Xformers: Efficient Attention for Image Classification
- Transformer Acceleration with Dynamic Sparse Attention
- UFO-ViT: High Performance Linear Vision Transformer without Softmax
- VQ-GNN: A Universal Framework to Scale up Graph Neural Networks using Vector Quantization
- The Sensory Neuron as a Transformer: Permutation-Invariant Neural Networks for Reinforcement Learning
- Not All Attention Is All You Need
- Sparse Attentive Memory Network for Click-through Rate Prediction with Long Sequences
- LazyFormer: Self Attention with Lazy Update
- Scatterbrain: Unifying Sparse and Low-rank Attention Approximation
- OmniNet: Omnidirectional Representations from Transformers
- Automated essay scoring using efficient transformer-based language models
- Stable, Fast and Accurate: Kernelized Attention with Relative Positional Encoding
- Optimizing Inference Performance of Transformers on CPUs
- Not All Memories are Created Equal: Learning to Forget by Expiring
- What Context Features Can Transformer Language Models Use?
- FMMformer: Efficient and Flexible Transformer via Decomposed Near-field and Far-field Attention
- Unifying Instance and Panoptic Segmentation with Dynamic Rank-1 Convolutions
- Redesigning the Transformer Architecture with Insights from Multi-particle Dynamical Systems
- Transformed CNNs: recasting pre-trained convolutional layers with self-attention
- A Singular Value Perspective on Model Robustness
- Video Transformer for Deepfake Detection with Incremental Learning
- HR-NAS: Searching Efficient High-Resolution Neural Architectures with Lightweight Transformers
- Transformation Invariant Cancerous Tissue Classification Using Spatially Transformed DenseNet
- Astronomical image time series classification using CONVolutional attENTION (ConvEntion)
- Translational Equivariance in Kernelizable Attention
- Armour: Generalizable Compact Self-Attention for Vision Transformers
- EL-Attention: Memory Efficient Lossless Attention for Generation
- Smart Bird: Learnable Sparse Attention for Efficient and Effective Transformer
- Pre-trained Language Model based Ranking in Baidu Search
- Transformer-F: A Transformer network with effective methods for learning universal sentence representation
- PairConnect: A Compute-Efficient MLP Alternative to Attention
- H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for Sequences
- Do Long-Range Language Models Actually Use Long-Range Context?
- Enabling Lightweight Fine-tuning for Pre-trained Language Model Compression based on Matrix Product Operators
- Exceeding the Limits of Visual-Linguistic Multi-Task Learning
- Updater-Extractor Architecture for Inductive World State Representations
- Linear Self-Attention Approximation via Trainable Feedforward Kernel
- Poly-NL: Linear Complexity Non-local Layers with Polynomials
- Efficient Transformer for Direct Speech Translation
- Semantic Frame Forecast
- Knowledge Transfer by Discriminative Pre-training for Academic Performance Prediction
- Decoupled Transformer for Scalable Inference in Open-domain Question Answering
- PermuteFormer: Efficient Relative Position Encoding for Long Sequences
- MATE: Multi-view Attention for Table Transformer Efficiency
- THG: Transformer with Hyperbolic Geometry
- Skyformer: Remodel Self-Attention with Gaussian Kernel and Nyström Method
- PlueckerNet: Learn to Register 3D Line Reconstructions
- Transfer training from smaller language model