An Attention Free Transformer
arXiv:2105.14103
Abstract
We introduce Attention Free Transformer (AFT), an efficient variant of Transformers that eliminates the need for dot product self attention. In an AFT layer, the key and value are first combined with a set of learned position biases, the result of which is multiplied with the query in an element-wise fashion. This new operation has a memory complexity linear w.r.t. both the context size and the dimension of features, making it compatible to both large input and model sizes. We also introduce AFT-local and AFT-conv, two model variants that take advantage of the idea of locality and spatial weight sharing while maintaining global connectivity. We conduct extensive experiments on two autoregressive modeling tasks (CIFAR10 and Enwik8) as well as an image recognition task (ImageNet-1K classification). We show that AFT demonstrates competitive performance on all the benchmarks, while providing excellent efficiency at the same time.
References in corpus (16)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Decoupled Weight Decay Regularization
- Deep Residual Learning for Image Recognition
- Categorical Reparameterization with Gumbel-Softmax
- Linformer: Self-Attention with Linear Complexity
- Generating Long Sequences with Sparse Transformers
- Reformer: The Efficient Transformer
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- Stand-Alone Self-Attention in Vision Models
- Synthesizer: Rethinking Self-Attention in Transformer Models
- Random Feature Attention
- Rethinking Attention with Performers
- Sparse Sinkhorn Attention
- Axial-DeepLab: Stand-Alone Axial-Attention for Panoptic Segmentation
- Interlaced Sparse Self-Attention for Semantic Segmentation
- LambdaNetworks: Modeling Long-Range Interactions Without Attention
Cited by in corpus (5)
- Filter-enhanced MLP is All You Need for Sequential Recommendation
- Counting Varying Density Crowds Through Density Guided Adaptive Selection CNN and Transformer Estimation
- Local Information Assisted Attention-free Decoder for Audio Captioning
- PoNet: Pooling Network for Efficient Token Mixing in Long Sequences
- Dispatcher: A Message-Passing Approach To Language Modelling