Efficient Transformers: A Survey
arXiv:2009.06732
Abstract
Transformer model architectures have garnered immense interest lately due to their effectiveness across a range of domains like language, vision and reinforcement learning. In the field of natural language processing for example, Transformers have become an indispensable staple in the modern deep learning stack. Recently, a dizzying number of "X-former" models have been proposed - Reformer, Linformer, Performer, Longformer, to name a few - which improve upon the original Transformer architecture, many of which make improvements around computational and memory efficiency. With the aim of helping the avid researcher navigate this flurry, this paper characterizes a large and thoughtful selection of recent efficiency-flavored "X-former" models, providing an organized and comprehensive overview of existing work and models across multiple domains.
Version 2: 2022 edition
Cited by in corpus (18)
- A Survey of Human-in-the-loop for Machine Learning
- A General Survey on Attention Mechanisms in Deep Learning
- A novel time-frequency Transformer based on self-attention mechanism and its application in fault diagnosis of rolling bearings
- Video Transformers: A Survey
- Enable Deep Learning on Mobile Devices: Methods, Systems, and Applications
- A Practical Survey on Faster and Lighter Transformers
- Attention-based graph neural networks: a survey
- An Algorithm-Hardware Co-Optimized Framework for Accelerating N:M Sparse Transformers
- Video Joint Modelling Based on Hierarchical Transformer for Co-summarization
- HeBERT & HebEMO: a Hebrew BERT Model and a Tool for Polarity Analysis and Emotion Recognition
- Attention Meets Perturbations: Robust and Interpretable Attention with Adversarial Training
- Text Guide: Improving the quality of long text classification by a text selection method based on feature importance
- Challenges in Domain-Specific Abstractive Summarization and How to Overcome them
- Divide and Rule: Effective Pre-Training for Context-Aware Multi-Encoder Translation Models
- Fourier Transformer: Fast Long Range Modeling by Removing Sequence Redundancy with FFT Operator
- Time-Space Transformers for Video Panoptic Segmentation
- Composition, Attention, or Both?
- On Inductive Biases for Machine Learning in Data Constrained Settings