Spatial Transformer Networks
arXiv:1506.02025
Abstract
Convolutional Neural Networks define an exceptionally powerful class of models, but are still limited by the lack of ability to be spatially invariant to the input data in a computationally and parameter efficient manner. In this work we introduce a new learnable module, the Spatial Transformer, which explicitly allows the spatial manipulation of data within the network. This differentiable module can be inserted into existing convolutional architectures, giving neural networks the ability to actively spatially transform feature maps, conditional on the feature map itself, without any extra training supervision or modification to the optimisation process. We show that the use of spatial transformers results in models which learn invariance to translation, scale, rotation and more generic warping, resulting in state-of-the-art performance on several benchmarks, and for a number of classes of transformations.
References in corpus (12)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Show, Attend and Tell: Neural Image Caption Generation with Visual Attention
- Improving neural networks by preventing co-adaptation of feature detectors
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition
- Bird Species Categorization Using Pose Normalized Deep Convolutional Nets
- Deep Networks with Internal Selective Attention through Feedback Connections
- Attention for Fine-Grained Categorization
- Locally Scale-Invariant Convolutional Neural Networks
- Neural Activation Constellations: Unsupervised Part Model Discovery with Convolutional Networks
- Contextual Action Recognition with R*CNN
- Bilinear CNNs for Fine-grained Visual Recognition
Cited by in corpus (39)
- End-to-End Unsupervised Deformable Image Registration with a Convolutional Neural Network
- Query2Label: A Simple Transformer Way to Multi-Label Classification
- LPRNet: License Plate Recognition via Deep Neural Networks
- TransAttUnet: Multi-level Attention-guided U-Net with Transformer for Medical Image Segmentation
- DenseCap: Fully Convolutional Localization Networks for Dense Captioning
- Efficient inference in occlusion-aware generative models of images
- FutureMapping: The Computational Structure of Spatial AI Systems
- Traffic Sign Classification Using Deep Inception Based Convolutional Networks
- Classification of Periodic Variable Stars with Novel Cyclic-Permutation Invariant Neural Networks
- SAN: Scale-Aware Network for Semantic Segmentation of High-Resolution Aerial Images
- Saliency-based Sequential Image Attention with Multiset Prediction
- Adversarial Image Registration with Application for MR and TRUS Image Fusion
- Dynamic Capacity Networks
- What Happened to My Dog in That Network: Unraveling Top-down Generators in Convolutional Neural Networks
- Focus Your Distribution: Coarse-to-Fine Non-Contrastive Learning for Anomaly Detection and Localization
- DeepIrisNet2: Learning Deep-IrisCodes from Scratch for Segmentation-Robust Visible Wavelength and Near Infrared Iris Recognition
- Leveraging Unsupervised Image Registration for Discovery of Landmark Shape Descriptor
- A State-of-the-art Survey of Artificial Neural Networks for Whole-slide Image Analysis:from Popular Convolutional Neural Networks to Potential Visual Transformers
- Neural Pharmacodynamic State Space Modeling
- Self-Supervised Generative Adversarial Network for Depth Estimation in Laparoscopic Images
- Sampling Equivariant Self-attention Networks for Object Detection in Aerial Images
- Deformed2Self: Self-Supervised Denoising for Dynamic Medical Imaging
- GradNets: Dynamic Interpolation Between Neural Architectures
- Representation and Correlation Enhanced Encoder-Decoder Framework for Scene Text Recognition
- CI-Net: Contextual Information for Joint Semantic Segmentation and Depth Estimation
- I2C2W: Image-to-Character-to-Word Transformers for Accurate Scene Text Recognition
- GMAIR: Unsupervised Object Detection Based on Spatial Attention and Gaussian Mixture
- Learning scale-variant and scale-invariant features for deep image classification
- Rotation Invariant Deep CBIR
- Comparison of Neuronal Attention Models
- A Controller-Recognizer Framework: How necessary is recognition for control?
- CelebHair: A New Large-Scale Dataset for Hairstyle Recommendation based on CelebA
- Adopting Robustness and Optimality in Fitting and Learning
- Built-in Elastic Transformations for Improved Robustness
- The Quo Vadis submission at Traffic4cast 2019
- A Benchmark Comparison of Visual Place Recognition Techniques for Resource-Constrained Embedded Platforms
- Recurrent Attention Models with Object-centric Capsule Representation for Multi-object Recognition
- End-to-end Ultrasound Frame to Volume Registration
- Monocular Depth Estimation with Directional Consistency by Deep Networks