Synthesizer: Rethinking Self-Attention in Transformer Models
arXiv:2005.00743
Abstract
The dot product self-attention is known to be central and indispensable to state-of-the-art Transformer models. But is it really required? This paper investigates the true importance and contribution of the dot product-based self-attention mechanism on the performance of Transformer models. Via extensive experiments, we find that (1) random alignment matrices surprisingly perform quite competitively and (2) learning attention weights from token-token (query-key) interactions is useful but not that important after all. To this end, we propose \textsc{Synthesizer}, a model that learns synthetic attention weights without token-token interactions. In our experiments, we first show that simple Synthesizers achieve highly competitive performance when compared against vanilla Transformer models across a range of tasks, including machine translation, language modeling, text generation and GLUE/SuperGLUE benchmarks. When composed with dot product attention, we find that Synthesizers consistently outperform Transformers. Moreover, we conduct additional comparisons of Synthesizers against Dynamic Convolutions, showing that simple Random Synthesizer is not only faster but also improves perplexity by a relative . Finally, we show that simple factorized Synthesizers can outperform Linformers on encoding only tasks.
ICML 2021
References in corpus (5)
Cited by in corpus (53)
- Transformers in Vision: A Survey
- MLP-Mixer: An all-MLP Architecture for Vision
- A General Survey on Attention Mechanisms in Deep Learning
- Long Range Arena: A Benchmark for Efficient Transformers
- A Practical Survey on Faster and Lighter Transformers
- Compressing Large-Scale Transformer-Based Models: A Case Study on BERT
- Random Feature Attention
- Hopfield Networks is All You Need
- Multi-Head Attention: Collaborate Instead of Concatenate
- Attention is Not All You Need: Pure Attention Loses Rank Doubly Exponentially with Depth
- Scale Efficiently: Insights from Pre-training and Fine-tuning Transformers
- Long-Short Transformer: Efficient Transformers for Language and Vision
- DeLighT: Deep and Light-weight Transformer
- An Attention Free Transformer
- Refiner: Refining Self-attention for Vision Transformers
- Transformer-based End-to-End Speech Recognition with Local Dense Synthesizer Attention
- Pay Attention to MLPs
- Are Pre-trained Convolutions Better than Pre-trained Transformers?
- Learning Graph Structures with Transformer for Multivariate Time Series Anomaly Detection in IoT
- Cluster-Former: Clustering-based Sparse Transformer for Long-Range Dependency Encoding
- CoCon: A Self-Supervised Approach for Controlled Text Generation
- Convolution-enhanced Evolving Attention Networks
- Primer: Searching for Efficient Transformers for Language Modeling
- AMBERT: A Pre-trained Language Model with Multi-Grained Tokenization
- Not All Attention Is All You Need
- Phonetic Posteriorgrams based Many-to-Many Singing Voice Conversion via Adversarial Training
- Fourier Transformer: Fast Long Range Modeling by Removing Sequence Redundancy with FFT Operator
- The Benchmark Lottery
- Summarize, Outline, and Elaborate: Long-Text Generation via Hierarchical Supervision from Extractive Summaries
- Eigen Analysis of Self-Attention and its Reconstruction from Partial Computation
- Direct Feedback Alignment Scales to Modern Deep Learning Tasks and Architectures
- HyperGrid: Efficient Multi-Task Transformers with Grid-wise Decomposable Hyper Projections
- FMMformer: Efficient and Flexible Transformer via Decomposed Near-field and Far-field Attention
- Adding Recurrence to Pretrained Transformers for Improved Efficiency and Context Size
- Redesigning the Transformer Architecture with Insights from Multi-particle Dynamical Systems
- The Role of Global and Local Context in Named Entity Recognition
- Attention Mechanism with Energy-Friendly Operations
- Current Limitations of Language Models: What You Need is Retrieval
- On Learning the Transformer Kernel
- Selective Knowledge Distillation for Neural Machine Translation
- Token Pooling in Vision Transformers
- PointMixer: MLP-Mixer for Point Cloud Understanding
- PairConnect: A Compute-Efficient MLP Alternative to Attention
- Semantic Representation and Inference for NLP
- H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for Sequences
- T-EMDE: Sketching-based global similarity for cross-modal retrieval
- Efficient Transformer for Direct Speech Translation
- Not all parameters are born equal: Attention is mostly what you need
- The Efficiency Misnomer
- PermuteFormer: Efficient Relative Position Encoding for Long Sequences
- Skyformer: Remodel Self-Attention with Gaussian Kernel and Nyström Method
- Do We Really Need That Many Parameters In Transformer For Extractive Summarization? Discourse Can Help !
- Input-independent Attention Weights Are Expressive Enough: A Study of Attention in Self-supervised Audio Transformers