Axial Attention in Multidimensional Transformers
arXiv:1912.12180
Abstract
We propose Axial Transformers, a self-attention-based autoregressive model for images and other data organized as high dimensional tensors. Existing autoregressive models either suffer from excessively large computational resource requirements for high dimensional data, or make compromises in terms of distribution expressiveness or ease of implementation in order to decrease resource requirements. Our architecture, by contrast, maintains both full expressiveness over joint distributions over data and ease of implementation with standard deep learning frameworks, while requiring reasonable memory and computation and achieving state-of-the-art results on standard generative modeling benchmarks. Our models are based on axial attention, a simple generalization of self-attention that naturally aligns with the multiple dimensions of the tensors in both the encoding and the decoding settings. Notably the proposed structure of the layers allows for the vast majority of the context to be computed in parallel during decoding without introducing any independence assumptions. This semi-parallel structure goes a long way to making decoding from even a very large Axial Transformer broadly applicable. We demonstrate state-of-the-art results for the Axial Transformer on the ImageNet-32 and ImageNet-64 image benchmarks as well as on the BAIR Robotic Pushing video benchmark. We open source the implementation of Axial Transformers.
10 pages
References in corpus (4)
Cited by in corpus (72)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Transformers in Vision: A Survey
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- Is Space-Time Attention All You Need for Video Understanding?
- Skillful Precipitation Nowcasting using Deep Generative Models of Radar
- DeepViT: Towards Deeper Vision Transformer
- MetNet: A Neural Weather Model for Precipitation Forecasting
- HANet: A Hierarchical Attention Network for Change Detection With Bitemporal Very-High-Resolution Remote Sensing Images
- XCiT: Cross-Covariance Image Transformers
- Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers
- Long Range Arena: A Benchmark for Efficient Transformers
- CSWin-UNet: Transformer UNet with Cross-Shaped Windows for Medical Image Segmentation
- VideoGPT: Video Generation using VQ-VAE and Transformers
- Perceiver: General Perception with Iterative Attention
- SAINT: Improved Neural Networks for Tabular Data via Row Attention and Contrastive Pre-Training
- Random Feature Attention
- Jukebox: A Generative Model for Music
- CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows
- Sparse DETR: Efficient End-to-End Object Detection with Learnable Sparsity
- Beyond 512 Tokens: Siamese Multi-depth Transformer-based Hierarchical Encoder for Long-Form Document Matching
- Chasing Sparsity in Vision Transformers: An End-to-End Exploration
- Protein language models trained on multiple sequence alignments learn phylogenetic relationships
- Axial-DeepLab: Stand-Alone Axial-Attention for Panoptic Segmentation
- Long-Short Transformer: Efficient Transformers for Language and Vision
- Effective Training Strategies for Deep-learning-based Precipitation Nowcasting and Estimation
- LambdaNetworks: Modeling Long-Range Interactions Without Attention
- Self-Attention Between Datapoints: Going Beyond Individual Input-Output Pairs in Deep Learning
- Pyramid Medical Transformer for Medical Image Segmentation
- RockGPT: Reconstructing three-dimensional digital rocks from single two-dimensional slice from the perspective of video generation
- DeepLab2: A TensorFlow Library for Deep Labeling
- MaX-DeepLab: End-to-End Panoptic Segmentation with Mask Transformers
- 3D Axial-Attention for Lung Nodule Classification
- Weaving Attention U-net: A Novel Hybrid CNN and Attention-based Method for Organs-at-risk Segmentation in Head and Neck CT Images
- Combiner: Full Attention Transformer with Sparse Computation Cost
- P2AT: Pyramid Pooling Axial Transformer for Real-time Semantic Segmentation
- Adaptive Fourier Neural Operators: Efficient Token Mixers for Transformers
- Global Self-Attention Networks for Image Recognition
- Visual Parser: Representing Part-whole Hierarchies with Transformers
- MS-nowcasting: Operational Precipitation Nowcasting with Convolutional LSTMs at Microsoft Weather
- Generative Adversarial Transformers
- Axial multi-layer perceptron architecture for automatic segmentation of choroid plexus in multiple sclerosis
- Transformers predicting the future. Applying attention in next-frame and time series forecasting
- Deep multi-stations weather forecasting: explainable recurrent convolutional neural networks
- Fourier Transformer: Fast Long Range Modeling by Removing Sequence Redundancy with FFT Operator
- Mitigation of Spatial Nonstationarity with Vision Transformers
- Scene Transformer: A unified architecture for predicting multiple agent trajectories
- Stable, Fast and Accurate: Kernelized Attention with Relative Positional Encoding
- Scalable and Accurate Self-supervised Multimodal Representation Learning without Aligned Video and Text Data
- FMMformer: Efficient and Flexible Transformer via Decomposed Near-field and Far-field Attention
- Hybrid Generative-Contrastive Representation Learning
- KVT: k-NN Attention for Boosting Vision Transformers
- Transformer Compressed Sensing via Global Image Tokens
- Redesigning the Transformer Architecture with Insights from Multi-particle Dynamical Systems
- Monocular Road Planar Parallax Estimation
- Improved Transformer for High-Resolution GANs
- Transformation Invariant Cancerous Tissue Classification Using Spatially Transformed DenseNet
- Axial Residual Networks for CycleGAN-based Voice Conversion
- baller2vec++: A Look-Ahead Multi-Entity Transformer For Modeling Coordinated Agents
- Hybrid and Collaborative Passage Reranking
- Community Research Earth Digital Intelligence Twin (CREDIT)
- Channelized Axial Attention for Semantic Segmentation -- Considering Channel Relation within Spatial Attention for Semantic Segmentation
- On the Distribution, Sparsity, and Inference-time Quantization of Attention Values in Transformers
- Is Batch Norm unique? An empirical investigation and prescription to emulate the best properties of common normalizers without batch dependence
- H-Transformer-1D: Fast One-Dimensional Hierarchical Attention for Sequences
- Video-based Person Re-identification without Bells and Whistles
- DR-TANet: Dynamic Receptive Temporal Attention Network for Street Scene Change Detection
- CpT: Convolutional Point Transformer for 3D Point Cloud Processing
- Generation and Simulation of Yeast Microscopy Imagery with Deep Learning
- Searching for TrioNet: Combining Convolution with Local and Global Self-Attention
- Portmanteauing Features for Scene Text Recognition
- Ripple Attention for Visual Perception with Sub-quadratic Complexity
- DCT: Dynamic Compressive Transformer for Modeling Unbounded Sequence