A Practical Survey on Faster and Lighter Transformers
arXiv:2103.14636 · doi:10.1145/3586074
Abstract
Recurrent neural networks are effective models to process sequences. However, they are unable to learn long-term dependencies because of their inherent sequential nature. As a solution, Vaswani et al. introduced the Transformer, a model solely based on the attention mechanism that is able to relate any two positions of the input sequence, hence modelling arbitrary long dependencies. The Transformer has improved the state-of-the-art across numerous sequence modelling tasks. However, its effectiveness comes at the expense of a quadratic computational and memory complexity with respect to the sequence length, hindering its adoption. Fortunately, the deep learning community has always been interested in improving the models' efficiency, leading to a plethora of solutions such as parameter sharing, pruning, mixed-precision, and knowledge distillation. Recently, researchers have directly addressed the Transformer's limitation by designing lower-complexity alternatives such as the Longformer, Reformer, Linformer, and Performer. However, due to the wide range of solutions, it has become challenging for researchers and practitioners to determine which methods to apply in practice in order to meet the desired trade-off between capacity, computation, and memory. This survey addresses this issue by investigating popular approaches to make Transformers faster and lighter and by providing a comprehensive explanation of the methods' strengths, limitations, and underlying assumptions.
ACM Computing Surveys; 40 pages, 18 figures, 4 tables
References in corpus (20)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Distilling the Knowledge in a Neural Network
- Sequence to Sequence Learning with Neural Networks
- Neural Architecture Search with Reinforcement Learning
- MLP-Mixer: An all-MLP Architecture for Vision
- Linformer: Self-Attention with Linear Complexity
- Attention is not Explanation
- Generating Long Sequences with Sparse Transformers
- Reformer: The Efficient Transformer
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- The Evolved Transformer
- The Reversible Residual Network: Backpropagation Without Storing Activations
- Lite Transformer with Long-Short Range Attention
- Sparse Sinkhorn Attention
- Nyströmformer: A Nyström-Based Algorithm for Approximating Self-Attention
- Finding Fast Transformers: One-Shot Neural Architecture Search by Component Composition
- Reconciling Modern Deep Learning with Traditional Optimization Analyses: The Intrinsic Learning Rate
- Transformers with Competitive Ensembles of Independent Mechanisms
- Transformer on a Diet
- Transformer-based Online Speech Recognition with Decoder-end Adaptive Computation Steps
Cited by in corpus (12)
- Neural Natural Language Processing for Long Texts: A Survey on Classification and Summarization
- Deep Learning-based Techniques for Integrated Sensing and Communication Systems: State-of-the-Art, Challenges, and Opportunities
- Representation learning for neural population activity with Neural Data Transformers
- Hybrid Quantum Vision Transformers for Event Classification in High Energy Physics
- Efficient and Private Federated Learning with Partially Trainable Networks
- Conditional computation in neural networks: principles and research trends
- TaylorShift: Shifting the Complexity of Self-Attention from Squared to Linear (and Back) using Taylor-Softmax
- GroupBERT: Enhanced Transformer Architecture with Efficient Grouped Structures
- Memory and Knowledge Augmented Language Models for Inferring Salience in Long-Form Stories
- NiNformer: A Network in Network Transformer with Token Mixing Generated Gating Function
- The MERIT Dataset: Modelling and Efficiently Rendering Interpretable Transcripts
- Forecasting Self-Similar User Traffic Demand Using Transformers in LEO Satellite Networks