Reformer: The Efficient Transformer
arXiv:2001.04451
Abstract
Large Transformer models routinely achieve state-of-the-art results on a number of tasks but training these models can be prohibitively costly, especially on long sequences. We introduce two techniques to improve the efficiency of Transformers. For one, we replace dot-product attention by one that uses locality-sensitive hashing, changing its complexity from O() to O(), where is the length of the sequence. Furthermore, we use reversible residual layers instead of the standard residuals, which allows storing activations only once in the training process instead of times, where is the number of layers. The resulting model, the Reformer, performs on par with Transformer models while being much more memory-efficient and much faster on long sequences.
ICLR 2020
References in corpus (4)
Cited by in corpus (92)
- Linformer: Self-Attention with Linear Complexity
- Informer: Beyond Efficient Transformer for Long Sequence Time-Series Forecasting
- Long Range Arena: A Benchmark for Efficient Transformers
- Residual Attention U-Net for Automated Multi-Class Segmentation of COVID-19 Chest CT Images
- Random Feature Attention
- Fastformer: Additive Attention Can Be All You Need
- Sparse Sinkhorn Attention
- LoFTR: Detector-Free Local Feature Matching with Transformers
- Luna: Linear Unified Nested Attention
- Keeping Your Eye on the Ball: Trajectory Attention in Video Transformers
- LambdaNetworks: Modeling Long-Range Interactions Without Attention
- The Expando-Mono-Duo Design Pattern for Text Ranking with Pretrained Sequence-to-Sequence Models
- Container: Context Aggregation Network
- AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing
- Learnable Fourier Features for Multi-Dimensional Spatial Positional Encoding
- CCVS: Context-aware Controllable Video Synthesis
- Masked Language Modeling for Proteins via Linearly Scalable Long-Context Transformers
- Nyströmformer: A Nyström-Based Algorithm for Approximating Self-Attention
- Combiner: Full Attention Transformer with Sparse Computation Cost
- A Survey of Deep Learning Approaches for OCR and Document Understanding
- Multi-head or Single-head? An Empirical Comparison for Transformer Training
- Neural Language Generation: Formulation, Methods, and Evaluation
- Recurrent Quantum Neural Networks
- GMAT: Global Memory Augmentation for Transformers
- Centroid Transformers: Learning to Abstract with Attention
- Conformer-Kernel with Query Term Independence for Document Retrieval
- Efficient Attentions for Long Document Summarization
- UFO-ViT: High Performance Linear Vision Transformer without Softmax
- lamBERT: Language and Action Learning Using Multimodal BERT
- Evolving Attention with Residual Convolutions
- LiftFormer: 3D Human Pose Estimation using attention models
- Generating Images with Sparse Representations
- Co-BERT: A Context-Aware BERT Retrieval Model Incorporating Local and Query-specific Context
- Scatterbrain: Unifying Sparse and Low-rank Attention Approximation
- Stable, Fast and Accurate: Kernelized Attention with Relative Positional Encoding
- Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network Models
- Fast Transformers with Clustered Attention
- Multiresolution and Multimodal Speech Recognition with Transformers
- LAMPRET: Layout-Aware Multimodal PreTraining for Document Understanding
- Shifted Chunk Transformer for Spatio-Temporal Representational Learning
- Sub-Linear Memory: How to Make Performers SLiM
- On the Expressive Power of Self-Attention Matrices
- Weak-Attention Suppression For Transformer Based Speech Recognition
- Transformers for Limit Order Books
- Not All Memories are Created Equal: Learning to Forget by Expiring
- What Context Features Can Transformer Language Models Use?
- Which *BERT? A Survey Organizing Contextualized Encoders
- FMMformer: Efficient and Flexible Transformer via Decomposed Near-field and Far-field Attention
- Multiple Sclerosis Severity Classification From Clinical Text
- Adding Recurrence to Pretrained Transformers for Improved Efficiency and Context Size
- TensorCoder: Dimension-Wise Attention via Tensor Representation for Natural Language Modeling
- Redesigning the Transformer Architecture with Insights from Multi-particle Dynamical Systems
- A Sliding-Window Approach to Automatic Creation of Meeting Minutes
- A Unified Efficient Pyramid Transformer for Semantic Segmentation
- Multi-scale Transformer Language Models
- Staircase Attention for Recurrent Processing of Sequences
- IOT: Instance-wise Layer Reordering for Transformer Structures
- General-Purpose User Embeddings based on Mobile App Usage
- EL-Attention: Memory Efficient Lossless Attention for Generation
- Adaptive Semiparametric Language Models
- Sparse Factorization of Large Square Matrices
- Smart Bird: Learnable Sparse Attention for Efficient and Effective Transformer
- Shatter: An Efficient Transformer Encoder with Single-Headed Self-Attention and Relative Sequence Partitioning
- Armour: Generalizable Compact Self-Attention for Vision Transformers
- Self-supervised Answer Retrieval on Clinical Notes
- Independent Encoder for Deep Hierarchical Unsupervised Image-to-Image Translation
- PairConnect: A Compute-Efficient MLP Alternative to Attention
- Beyond Nyströmformer -- Approximation of self-attention by Spectral Shifting
- Feedback Attention for Cell Image Segmentation
- Do Long-Range Language Models Actually Use Long-Range Context?
- Learning Elastic Embeddings for Customizing On-Device Recommenders
- Token Shift Transformer for Video Classification
- H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for Sequences
- SHAPE: Shifted Absolute Position Embedding for Transformers
- Long-Range Modeling of Source Code Files with eWASH: Extended Window Access by Syntax Hierarchy
- Understanding and Overcoming the Challenges of Efficient Transformer Quantization
- Leveraging redundancy in attention with Reuse Transformers
- HETFORMER: Heterogeneous Transformer with Sparse Attention for Long-Text Extractive Summarization
- Skyformer: Remodel Self-Attention with Gaussian Kernel and Nyström Method
- A light transformer for speech-to-intent applications
- Reciprocal Supervised Learning Improves Neural Machine Translation
- Does Dialog Length matter for Next Response Selection task? An Empirical Study
- Relation/Entity-Centric Reading Comprehension
- PermuteFormer: Efficient Relative Position Encoding for Long Sequences
- Graph Conditioned Sparse-Attention for Improved Source Code Understanding
- Plot-guided Adversarial Example Construction for Evaluating Open-domain Story Generation
- MATE: Multi-view Attention for Table Transformer Efficiency
- A Strong Baseline for Query Efficient Attacks in a Black Box Setting
- Knowledge Transfer by Discriminative Pre-training for Academic Performance Prediction
- Efficient Transformer for Direct Speech Translation
- Can Transformer Models Measure Coherence In Text? Re-Thinking the Shuffle Test
- No Answer is Better Than Wrong Answer: A Reflection Model for Document Level Machine Reading Comprehension