Rethinking Attention with Performers
arXiv:2009.14794
Abstract
We introduce Performers, Transformer architectures which can estimate regular (softmax) full-rank-attention Transformers with provable accuracy, but using only linear (as opposed to quadratic) space and time complexity, without relying on any priors such as sparsity or low-rankness. To approximate softmax attention-kernels, Performers use a novel Fast Attention Via positive Orthogonal Random features approach (FAVOR+), which may be of independent interest for scalable kernel methods. FAVOR+ can be also used to efficiently model kernelizable attention mechanisms beyond softmax. This representational power is crucial to accurately compare softmax with other kernels for the first time on large-scale tasks, beyond the reach of regular Transformers, and investigate optimal attention-kernels. Performers are linear architectures fully compatible with regular Transformers and with strong theoretical guarantees: unbiased or nearly-unbiased estimation of the attention matrix, uniform convergence and low estimation variance. We tested Performers on a rich set of tasks stretching from pixel-prediction through text models to protein sequence modeling. We demonstrate competitive results with other examined efficient sparse and dense attention methods, showcasing effectiveness of the novel attention-learning paradigm leveraged by Performers.
Published as a conference paper + oral presentation at ICLR 2021. 38 pages. See https://github.com/google-research/google-research/tree/master/protein_lm for protein language model code, and https://github.com/google-research/google-research/tree/master/performer for Performer code. See https://ai.googleblog.com/2020/10/rethinking-attention-with-performers.html for Google AI Blog
References in corpus (8)
- Linformer: Self-Attention with Linear Complexity
- Generating Long Sequences with Sparse Transformers
- Conformer: Convolution-augmented Transformer for Speech Recognition
- Reformer: The Efficient Transformer
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- Orthogonal Random Features
- Orthogonal Estimation of Wasserstein Distances
- Demystifying Orthogonal Monte Carlo and Beyond
Cited by in corpus (85)
- On the Opportunities and Risks of Foundation Models
- CoAtNet: Marrying Convolution and Attention for All Data Sizes
- SpookyNet: Learning Force Fields with Electronic Degrees of Freedom and Nonlocal Effects
- LocalViT: Analyzing Locality in Vision Transformers
- TransGAN: Two Pure Transformers Can Make One Strong GAN, and That Can Scale Up
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- XCiT: Cross-Covariance Image Transformers
- Long Range Arena: A Benchmark for Efficient Transformers
- The NLP Cookbook: Modern Recipes for Transformer based Deep Learning Architectures
- CANINE: Pre-training an Efficient Tokenization-Free Encoder for Language Representation
- Random Feature Attention
- End-to-End Object Detection with Adaptive Clustering Transformer
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows
- Chasing Sparsity in Vision Transformers: An End-to-End Exploration
- Charformer: Fast Character Transformers via Gradient-based Subword Tokenization
- Long-Range Transformers for Dynamic Spatiotemporal Forecasting
- Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers
- Multiscale Vision Transformers
- Luna: Linear Unified Nested Attention
- LambdaNetworks: Modeling Long-Range Interactions Without Attention
- An Attention Free Transformer
- HTLM: Hyper-Text Pre-Training and Prompting of Language Models
- Unsupervised Brain Anomaly Detection and Segmentation with Transformers
- Automated Identification of Cell Populations in Flow Cytometry Data with Transformers
- Nyströmformer: A Nyström-Based Algorithm for Approximating Self-Attention
- Text Guide: Improving the quality of long text classification by a text selection method based on feature importance
- Adaptive Fourier Neural Operators: Efficient Token Mixers for Transformers
- A Survey of Deep Learning Approaches for OCR and Document Understanding
- Interpreting Deep Learning Models in Natural Language Processing: A Review
- SLAPS: Self-Supervision Improves Structure Learning for Graph Neural Networks
- Relative Positional Encoding for Transformers with Linear Complexity
- Kaleido-BERT: Vision-Language Pre-training on Fashion Domain
- Hi-BEHRT: Hierarchical Transformer-based model for accurate prediction of clinical events using multimodal longitudinal electronic health records
- UFO-ViT: High Performance Linear Vision Transformer without Softmax
- Efficient conformer-based speech recognition with linear attention
- Efficient and Private Federated Learning with Partially Trainable Networks
- The Sensory Neuron as a Transformer: Permutation-Invariant Neural Networks for Reinforcement Learning
- LazyFormer: Self Attention with Lazy Update
- OmniNet: Omnidirectional Representations from Transformers
- Length-Adaptive Transformer: Train Once with Length Drop, Use Anytime with Search
- Scatterbrain: Unifying Sparse and Low-rank Attention Approximation
- 3D Object Detection with Pointformer
- Noise-Robust End-to-End Quantum Control using Deep Autoregressive Policy Networks
- ChineseBERT: Chinese Pretraining Enhanced by Glyph and Pinyin Information
- Automated essay scoring using efficient transformer-based language models
- Rethinking Lifelong Sequential Recommendation with Incremental Multi-Interest Attention
- RomeBERT: Robust Training of Multi-Exit BERT
- Attention Approximates Sparse Distributed Memory
- Eigen Analysis of Self-Attention and its Reconstruction from Partial Computation
- Sub-Linear Memory: How to Make Performers SLiM
- Group Equivariant Stand-Alone Self-Attention For Vision
- Redesigning the Transformer Architecture with Insights from Multi-particle Dynamical Systems
- PnP-DETR: Towards Efficient Visual Analysis with Transformers
- Unifying Instance and Panoptic Segmentation with Dynamic Rank-1 Convolutions
- A Singular Value Perspective on Model Robustness
- Transformed CNNs: recasting pre-trained convolutional layers with self-attention
- Learning Self-Modulating Attention in Continuous Time Space with Applications to Sequential Recommendation
- VEGN: Variant Effect Prediction with Graph Neural Networks
- Sparse Factorization of Large Square Matrices
- Sparse Attention with Linear Units
- TED-net: Convolution-free T2T Vision Transformer-based Encoder-decoder Dilation network for Low-dose CT Denoising
- Armour: Generalizable Compact Self-Attention for Vision Transformers
- FNetAR: Mixing Tokens with Autoregressive Fourier Transforms
- Updater-Extractor Architecture for Inductive World State Representations
- Token Shift Transformer for Video Classification
- Pre-trained Language Model based Ranking in Baidu Search
- PairConnect: A Compute-Efficient MLP Alternative to Attention
- Compositional Attention: Disentangling Search and Retrieval
- T-EMDE: Sketching-based global similarity for cross-modal retrieval
- Audiomer: A Convolutional Transformer For Keyword Spotting
- Ripple Attention for Visual Perception with Sub-quadratic Complexity
- A Survey on Green Deep Learning
- Kernel Deformed Exponential Families for Sparse Continuous Attention
- Graph Conditioned Sparse-Attention for Improved Source Code Understanding
- Detect the Interactions that Matter in Matter: Geometric Attention for Many-Body Systems
- SHORING: Design Provable Conditional High-Order Interaction Network via Symbolic Testing
- Grid Partitioned Attention: Efficient TransformerApproximation with Inductive Bias for High Resolution Detail Generation
- Knowledge Transfer by Discriminative Pre-training for Academic Performance Prediction
- Transfer training from smaller language model
- Understanding and Overcoming the Challenges of Efficient Transformer Quantization
- DCT: Dynamic Compressive Transformer for Modeling Unbounded Sequence
- Leveraging redundancy in attention with Reuse Transformers
- Skyformer: Remodel Self-Attention with Gaussian Kernel and Nyström Method
- PlueckerNet: Learn to Register 3D Line Reconstructions
- Does Dialog Length matter for Next Response Selection task? An Empirical Study