Ripple Attention for Visual Perception with Sub-quadratic Complexity
arXiv:2110.02453
Abstract
Transformer architectures are now central to sequence modeling tasks. At its heart is the attention mechanism, which enables effective modeling of long-term dependencies in a sequence. Recently, transformers have been successfully applied in the computer vision domain, where 2D images are first segmented into patches and then treated as 1D sequences. Such linearization, however, impairs the notion of spatial locality in images, which bears important visual clues. To bridge the gap, we propose ripple attention, a sub-quadratic attention mechanism for vision transformers. Built upon the recent kernel-based efficient attention mechanisms, we design a novel dynamic programming algorithm that weights contributions of different tokens to a query with respect to their relative spatial distances in the 2D space in linear observed time. Extensive experiments and analyses demonstrate the effectiveness of ripple attention on various visual tasks.
19 pages, 2 figures, ICML 2022 camera ready
References in corpus (24)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- PVT v2: Improved Baselines with Pyramid Vision Transformer
- Is Space-Time Attention All You Need for Video Understanding?
- Linformer: Self-Attention with Linear Complexity
- Generating Long Sequences with Sparse Transformers
- Axial Attention in Multidimensional Transformers
- Early Convolutions Help Transformers See Better
- Reformer: The Efficient Transformer
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- LocalViT: Analyzing Locality in Vision Transformers
- Focal Self-attention for Local-Global Interactions in Vision Transformers
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Long Range Arena: A Benchmark for Efficient Transformers
- CvT: Introducing Convolutions to Vision Transformers
- ViTAE: Vision Transformer Advanced by Exploring Intrinsic Inductive Bias
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows
- How Much Position Information Do Convolutional Neural Networks Encode?
- Rethinking Spatial Dimensions of Vision Transformers
- ConTNet: Why not use convolution and transformer at the same time?
- Video Swin Transformer
- MultiGrain: a unified image embedding for classes and instances
- Linear Transformers Are Secretly Fast Weight Programmers
- Relative Positional Encoding for Transformers with Linear Complexity
- Stable, Fast and Accurate: Kernelized Attention with Relative Positional Encoding