Residual Attention Network for Image Classification
arXiv:1704.06904
Abstract
In this work, we propose "Residual Attention Network", a convolutional neural network using attention mechanism which can incorporate with state-of-art feed forward network architecture in an end-to-end training fashion. Our Residual Attention Network is built by stacking Attention Modules which generate attention-aware features. The attention-aware features from different modules change adaptively as layers going deeper. Inside each Attention Module, bottom-up top-down feedforward structure is used to unfold the feedforward and feedback attention process into a single feedforward process. Importantly, we propose attention residual learning to train very deep Residual Attention Networks which can be easily scaled up to hundreds of layers. Extensive analyses are conducted on CIFAR-10 and CIFAR-100 datasets to verify the effectiveness of every module mentioned above. Our Residual Attention Network achieves state-of-the-art object recognition performance on three benchmark datasets including CIFAR-10 (3.90% error), CIFAR-100 (20.45% error) and ImageNet (4.8% single model and single crop, top-5 error). Note that, our method achieves 0.6% top-1 accuracy improvement with 46% trunk depth and 69% forward FLOPs comparing to ResNet-200. The experiment also demonstrates that our network is robust against noisy labels.
accepted to CVPR2017
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Wide Residual Networks
- Going Deeper with Convolutions
- Recurrent Models of Visual Attention
- DRAW: A Recurrent Neural Network For Image Generation
- Fully Convolutional Networks for Semantic Segmentation
- Learning Deconvolution Network for Semantic Segmentation
- Training Convolutional Networks with Noisy Labels
- Aggregated Residual Transformations for Deep Neural Networks
- Multimodal Residual Learning for Visual QA
- Deep Networks with Internal Selective Attention through Feedback Connections
Cited by in corpus (55)
- ABCNet: Attentive Bilateral Contextual Network for Efficient Semantic Segmentation of Fine-Resolution Remote Sensing Images
- CBAM: Convolutional Block Attention Module
- An Explainable 3D Residual Self-Attention Deep Neural Network FOR Joint Atrophy Localization and Alzheimer's Disease Diagnosis using Structural MRI
- Transparency by Design: Closing the Gap Between Performance and Interpretability in Visual Reasoning
- StainNet: a fast and robust stain normalization network
- Fixed Pattern Noise Reduction for Infrared Images Based on Cascade Residual Attention CNN
- Selective Kernel Networks
- Cross-regional oil palm tree counting and detection via multi-level attention domain adaptation network
- Deep Metric Learning-based Image Retrieval System for Chest Radiograph and its Clinical Applications in COVID-19
- Saliency for Fine-grained Object Recognition in Domains with Scarce Training Data
- Reading Scene Text with Attention Convolutional Sequence Modeling
- A Survey on Deep Learning Methods for Robot Vision
- Support Vector Guided Softmax Loss for Face Recognition
- Dynamic Neural Networks: A Survey
- Interaction-and-Aggregation Network for Person Re-identification
- Prior-Knowledge and Attention-based Meta-Learning for Few-Shot Learning
- PVNet: A Joint Convolutional Network of Point Cloud and Multi-View for 3D Shape Recognition
- One-shot Face Reenactment
- WIDER Face and Pedestrian Challenge 2018: Methods and Results
- Improving Object Detection from Scratch via Gated Feature Reuse
- Towards Interpretable Reinforcement Learning Using Attention Augmented Agents
- End-to-End Multi-Task Learning with Attention
- Cardiac Segmentation on CT Images through Shape-Aware Contour Attentions
- SegVoxelNet: Exploring Semantic Context and Depth-aware Features for 3D Vehicle Detection from Point Cloud
- Deep attention-based classification network for robust depth prediction
- Micro-Attention for Micro-Expression recognition
- FaceX-Zoo: A PyTorch Toolbox for Face Recognition
- Mis-classified Vector Guided Softmax Loss for Face Recognition
- Few-shot Classification via Adaptive Attention
- Domain Attention Consistency for Multi-Source Domain Adaptation
- DensSiam: End-to-End Densely-Siamese Network with Self-Attention Model for Object Tracking
- Towards Visually Explaining Similarity Models
- Residual Convolutional Neural Network Revisited with Active Weighted Mapping
- MAANet: Multi-view Aware Attention Networks for Image Super-Resolution
- Channel Locality Block: A Variant of Squeeze-and-Excitation
- Classification-driven Single Image Dehazing
- Attention Mechanisms for Object Recognition with Event-Based Cameras
- Deep Discriminative Representation Learning with Attention Map for Scene Classification
- Attentive Filtering Networks for Audio Replay Attack Detection
- Attention Guided Metal Artifact Correction in MRI using Deep Neural Networks
- Dynamic Filtering with Large Sampling Field for ConvNets
- Multiple Attentional Pyramid Networks for Chinese Herbal Recognition
- Spatial-Aware Non-Local Attention for Fashion Landmark Detection
- Learning To Pay Attention To Mistakes
- Action Recognition with Spatio-Temporal Visual Attention on Skeleton Image Sequences
- FineFool: Fine Object Contour Attack via Attention
- Towards Understanding the Effectiveness of Attention Mechanism
- The Devils in the Point Clouds: Studying the Robustness of Point Cloud Convolutions
- Gated Texture CNN for Efficient and Configurable Image Denoising
- Representation based and Attention augmented Meta learning
- Does Haze Removal Help CNN-based Image Classification?
- An Attention Model for group-level emotion recognition
- Pixel-Semantic Revise of Position Learning A One-Stage Object Detector with A Shared Encoder-Decoder
- Super Interaction Neural Network
- Selective Information Passing for MR/CT Image Segmentation