Squeeze-and-Excitation Networks
arXiv:1709.01507
Abstract
The central building block of convolutional neural networks (CNNs) is the convolution operator, which enables networks to construct informative features by fusing both spatial and channel-wise information within local receptive fields at each layer. A broad range of prior research has investigated the spatial component of this relationship, seeking to strengthen the representational power of a CNN by enhancing the quality of spatial encodings throughout its feature hierarchy. In this work, we focus instead on the channel relationship and propose a novel architectural unit, which we term the "Squeeze-and-Excitation" (SE) block, that adaptively recalibrates channel-wise feature responses by explicitly modelling interdependencies between channels. We show that these blocks can be stacked together to form SENet architectures that generalise extremely effectively across different datasets. We further demonstrate that SE blocks bring significant improvements in performance for existing state-of-the-art CNNs at slight additional computational cost. Squeeze-and-Excitation Networks formed the foundation of our ILSVRC 2017 classification submission which won first place and reduced the top-5 error to 2.251%, surpassing the winning entry of 2016 by a relative improvement of ~25%. Models and code are available at https://github.com/hujie-frank/SENet.
journal version of the CVPR 2018 paper, accepted by TPAMI
Cited by in corpus (242)
- Res2Net: A New Multi-scale Backbone Architecture
- Object Detection in Optical Remote Sensing Images: A Survey and A New Benchmark
- Efficient Multi-Scale Attention Module with Cross-Spatial Learning
- ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification
- Multivariate LSTM-FCNs for Time Series Classification
- UIU-Net: U-Net in U-Net for Infrared Small Object Detection
- fastai: A Layered API for Deep Learning
- Deep Face Recognition: A Survey
- Remote Sensing Image Scene Classification Meets Deep Learning: Challenges, Methods, Benchmarks, and Opportunities
- ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices
- Medical Image Segmentation Using Deep Learning: A Survey
- Deep neural network models for computational histopathology: A survey
- BACH: Grand Challenge on Breast Cancer Histology Images
- Human Action Recognition from Various Data Modalities: A Review
- AnatomyNet: Deep Learning for Fast and Fully Automated Whole-volume Segmentation of Head and Neck Anatomy
- Wireless Image Transmission Using Deep Source Channel Coding With Attention Modules
- The Devil is in the Channels: Mutual-Channel Loss for Fine-Grained Image Classification
- Avoiding Overfitting: A Survey on Regularization Methods for Convolutional Neural Networks
- Real-Time Polyp Detection, Localization and Segmentation in Colonoscopy Using Deep Learning
- Advancements in Image Classification using Convolutional Neural Network
- Path Aggregation Network for Instance Segmentation
- Multi-stage Attention ResU-Net for Semantic Segmentation of Fine-Resolution Remote Sensing Images
- Action Transformer: A Self-Attention Model for Short-Time Pose-Based Human Action Recognition
- Modality specific U-Net variants for biomedical image segmentation: A survey
- ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design
- Global Guidance Network for Breast Lesion Segmentation in Ultrasound Images
- Compute and Energy Consumption Trends in Deep Learning Inference
- A Comprehensive Review for Breast Histopathology Image Analysis Using Classical and Deep Neural Networks
- Insights into LSTM Fully Convolutional Networks for Time Series Classification
- MASTER: Multi-Aspect Non-local Network for Scene Text Recognition
- VehicleNet: Learning Robust Visual Representation for Vehicle Re-identification
- Deep learning at the shallow end: Malware classification for non-domain experts
- Causal Attention for Interpretable and Generalizable Graph Classification
- On Extended Long Short-term Memory and Dependent Bidirectional Recurrent Neural Network
- RMDL: Recalibrated multi-instance deep learning for whole slide gastric image classification
- Skin Lesion Classification Using CNNs with Patch-Based Attention and Diagnosis-Guided Loss Weighting
- Anomaly Detection-Inspired Few-Shot Medical Image Segmentation Through Self-Supervision With Supervoxels
- From Handcrafted to Deep Features for Pedestrian Detection: A Survey
- Dynamic Feature Integration for Simultaneous Detection of Salient Object, Edge and Skeleton
- Looking Outside the Window: Wide-Context Transformer for the Semantic Segmentation of High-Resolution Remote Sensing Images
- StegNet: Mega Image Steganography Capacity with Deep Convolutional Network
- Road Segmentation for Remote Sensing Images using Adversarial Spatial Pyramid Networks
- A3CLNN: Spatial, Spectral and Multiscale Attention ConvLSTM Neural Network for Multisource Remote Sensing Data Classification
- Improving the Harmony of the Composite Image by Spatial-Separated Attention Module
- Is Someone Speaking? Exploring Long-term Temporal Features for Audio-visual Active Speaker Detection
- Learning Hierarchical Attention for Weakly-supervised Chest X-Ray Abnormality Localization and Diagnosis
- BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation
- Salient Object Detection in Optical Remote Sensing Images Driven by Transformer
- Weakly Supervised Deep Learning for Thoracic Disease Classification and Localization on Chest X-rays
- Federated Continual Learning via Knowledge Fusion: A Survey
- Siamese Attentional Keypoint Network for High Performance Visual Tracking
- Multi-task Learning with Coarse Priors for Robust Part-aware Person Re-identification
- FoodAI: Food Image Recognition via Deep Learning for Smart Food Logging
- Decoding and mapping task states of the human brain via deep learning
- Learning Reinforced Attentional Representation for End-to-End Visual Tracking
- Machine learning with neural networks
- Rethinking and Designing a High-performing Automatic License Plate Recognition Approach
- Hierarchical Paired Channel Fusion Network for Street Scene Change Detection
- ECAPA-TDNN Embeddings for Speaker Diarization
- FocusNetv2: Imbalanced Large and Small Organ Segmentation with Adversarial Shape Constraint for Head and Neck CT Images
- Optimized Deep Encoder-Decoder Methods for Crack Segmentation
- Attention-based Pyramid Aggregation Network for Visual Place Recognition
- LO-Det: Lightweight Oriented Object Detection in Remote Sensing Images
- Deep Polynomial Neural Networks
- ADAM Challenge: Detecting Age-related Macular Degeneration from Fundus Images
- Sequential vessel segmentation via deep channel attention network
- Robust Ultra-wideband Range Error Mitigation with Deep Learning at the Edge
- Parameter-Efficient Person Re-identification in the 3D Space
- Sound source detection, localization and classification using consecutive ensemble of CRNN models
- Generalized Domain Conditioned Adaptation Network
- Meta Balanced Network for Fair Face Recognition
- Stage-Aware Feature Alignment Network for Real-Time Semantic Segmentation of Street Scenes
- ECOVNet: An Ensemble of Deep Convolutional Neural Networks Based on EfficientNet to Detect COVID-19 From Chest X-rays
- Semi-Heterogeneous Three-Way Joint Embedding Network for Sketch-Based Image Retrieval
- Integrating Frequency Translational Invariance in TDNNs and Frequency Positional Information in 2D ResNets to Enhance Speaker Verification
- On Deep Learning Techniques to Boost Monocular Depth Estimation for Autonomous Navigation
- Convolutional Neural Networks with Gated Recurrent Connections
- RangeSeg: Range-Aware Real Time Segmentation of 3D LiDAR Point Clouds
- The IDLAB VoxSRC-20 Submission: Large Margin Fine-Tuning and Quality-Aware Score Calibration in DNN Based Speaker Verification
- HEU Emotion: A Large-scale Database for Multi-modal Emotion Recognition in the Wild
- Difference-Based Deep Learning Framework for Stress Predictions in Heterogeneous Media
- A Unified Framework for Generalized Low-Shot Medical Image Segmentation with Scarce Data
- Compacting Deep Neural Networks for Internet of Things: Methods and Applications
- AeroRIT: A New Scene for Hyperspectral Image Analysis
- HEROHE Challenge: assessing HER2 status in breast cancer without immunohistochemistry or in situ hybridization
- Dynamic Instance Domain Adaptation
- CovTANet: A Hybrid Tri-level Attention Based Network for Lesion Segmentation, Diagnosis, and Severity Prediction of COVID-19 Chest CT Scans
- Bi-directional Cross-Modality Feature Propagation with Separation-and-Aggregation Gate for RGB-D Semantic Segmentation
- A Efficient Multimodal Framework for Large Scale Emotion Recognition by Fusing Music and Electrodermal Activity Signals
- Video Semantic Segmentation with Distortion-Aware Feature Correction
- Relational Deep Feature Learning for Heterogeneous Face Recognition
- Neutrino interaction classification with a convolutional neural network in the DUNE far detector
- Fine-Grained Fashion Similarity Prediction by Attribute-Specific Embedding Learning
- A Two-Stream Symmetric Network with Bidirectional Ensemble for Aerial Image Matching
- Depthwise Separable Convolutional ResNet with Squeeze-and-Excitation Blocks for Small-footprint Keyword Spotting
- Domain Adaptation Meets Zero-Shot Learning: An Annotation-Efficient Approach to Multi-Modality Medical Image Segmentation
- A Two-Stage Cascade Model with Variational Autoencoders and Attention Gates for MRI Brain Tumor Segmentation
- AugFPN: Improving Multi-scale Feature Learning for Object Detection
- A Comprehensive Review of Modern Object Segmentation Approaches
- Attention guided global enhancement and local refinement network for semantic segmentation
- Distortion-aware Monocular Depth Estimation for Omnidirectional Images
- Underwater Acoustic Target Recognition based on Smoothness-inducing Regularization and Spectrogram-based Data Augmentation
- Fingerprint Presentation Attack Detection by Channel-wise Feature Denoising
- Predicting Elastic Properties of Materials from Electronic Charge Density Using 3D Deep Convolutional Neural Networks
- Acoustic Scene Classification with Squeeze-Excitation Residual Networks
- Danish Fungi 2020 -- Not Just Another Image Recognition Dataset
- Contrastive Self-supervised Neural Architecture Search
- Comparative evaluation of CNN architectures for Image Caption Generation
- General audio tagging with ensembling convolutional neural network and statistical features
- Feature Calibration Network for Occluded Pedestrian Detection
- Calibration of Deep Probabilistic Models with Decoupled Bayesian Neural Networks
- PSLT: A Light-weight Vision Transformer with Ladder Self-Attention and Progressive Shift
- AP-MTL: Attention Pruned Multi-task Learning Model for Real-time Instrument Detection and Segmentation in Robot-assisted Surgery
- Towards Better Accuracy-efficiency Trade-offs: Divide and Co-training
- Category-Specific CNN for Visual-aware CTR Prediction at JD.com
- COVID-MTL: Multitask Learning with Shift3D and Random-weighted Loss for Automated Diagnosis and Severity Assessment of COVID-19
- CHS-Net: A Deep learning approach for hierarchical segmentation of COVID-19 infected CT images
- White Box Methods for Explanations of Convolutional Neural Networks in Image Classification Tasks
- Ultrasound Image Representation Learning by Modeling Sonographer Visual Attention
- Embedded Self-Distillation in Compact Multi-Branch Ensemble Network for Remote Sensing Scene Classification
- Bio-Inspired Representation Learning for Visual Attention Prediction
- An Initial Investigation for Detecting Vocoder Fingerprints of Fake Audio
- Estimating Parameters of the Tree Root in Heterogeneous Soil Environments via Mask-Guided Multi-Polarimetric Integration Neural Network
- A novel classification-selection approach for the self updating of template-based face recognition systems
- Sejong Face Database: A Multi-Modal Disguise Face Database
- EOCSA: Predicting Prognosis of Epithelial Ovarian Cancer with Whole Slide Histopathological Images
- Interpretable deep learning for nuclear deformation in heavy ion collisions
- Multi-modal land cover mapping of remote sensing images using pyramid attention and gated fusion networks
- Equalization Loss for Long-Tailed Object Recognition
- Style Mixer: Semantic-aware Multi-Style Transfer Network
- Regular Polytope Networks
- Deep Learning for Pancreas Segmentation: a Systematic Review
- Enhanced Standard Compatible Image Compression Framework based on Auxiliary Codec Networks
- PixelGame: Infrared small target segmentation as a Nash equilibrium
- ORDNet: Capturing Omni-Range Dependencies for Scene Parsing
- EfficientTDNN: Efficient Architecture Search for Speaker Recognition
- Efficient Spiking Neural Networks with Logarithmic Temporal Coding
- Sequential Image-based Attention Network for Inferring Force Estimation without Haptic Sensor
- Spatio-Temporal FAST 3D Convolutions for Human Action Recognition
- Towards Visual Distortion in Black-Box Attacks
- Lightweight Residual Densely Connected Convolutional Neural Network
- Distribution-aware Margin Calibration for Semantic Segmentation in Images
- Geometric Approaches to Increase the Expressivity of Deep Neural Networks for MR Reconstruction
- Side-Aware Boundary Localization for More Precise Object Detection
- Modality Attention and Sampling Enables Deep Learning with Heterogeneous Marker Combinations in Fluorescence Microscopy
- Cross-Lingual Speaker Verification with Domain-Balanced Hard Prototype Mining and Language-Dependent Score Normalization
- Anchor Pruning for Object Detection
- MERANet: Facial Micro-Expression Recognition using 3D Residual Attention Network
- Learning Models of Individual Behavior in Chess
- Replay and Synthetic Speech Detection with Res2net Architecture
- Improved YOLOv5s model for key components detection of power transmission lines
- A Dataset and Benchmark Towards Multi-Modal Face Anti-Spoofing Under Surveillance Scenarios
- Multitask Balanced and Recalibrated Network for Medical Code Prediction
- Pushing the Limits of Non-Autoregressive Speech Recognition
- A New Benchmark and Model for Challenging Image Manipulation Detection
- HOPE: Hybrid-granularity Ordinal Prototype Learning for Progression Prediction of Mild Cognitive Impairment
- Pollen Grain Microscopic Image Classification Using an Ensemble of Fine-Tuned Deep Convolutional Neural Networks
- Multi-Scale Deep Learning for Estimating Horizontal Velocity Fields on the Solar Surface
- Analysis of an adaptive lead weighted ResNet for multiclass classification of 12-lead ECGs
- Pairwise Comparison Network for Remote Sensing Scene Classification
- OnDev-LCT: On-Device Lightweight Convolutional Transformers towards federated learning
- Efficient Folded Attention for 3D Medical Image Reconstruction and Segmentation
- Segmentation-free PVC for Cardiac SPECT using a Densely-connected Multi-dimensional Dynamic Network
- Dynamic Refinement Network for Oriented and Densely Packed Object Detection
- Results of the NeurIPS'21 Challenge on Billion-Scale Approximate Nearest Neighbor Search
- Attention Convolutional Binary Neural Tree for Fine-Grained Visual Categorization
- Lightweight Stepless Super-Resolution of Remote Sensing Images via Saliency-Aware Dynamic Routing Strategy
- Adversarial Margin Maximization Networks
- How Will It Drape Like? Capturing Fabric Mechanics from Depth Images
- Understanding the computational demands underlying visual reasoning
- Learning to play the Chess Variant Crazyhouse above World Champion Level with Deep Neural Networks and Human Data
- A deep learning-based framework for segmenting invisible clinical target volumes with estimated uncertainties for post-operative prostate cancer radiotherapy
- Multitask Recalibrated Aggregation Network for Medical Code Prediction
- Learning Motion and Temporal Cues for Unsupervised Video Object Segmentation
- On the impact of selected modern deep-learning techniques to the performance and celerity of classification models in an experimental high-energy physics use case
- p-Meta: Towards On-device Deep Model Adaptation
- Tensor Low-Rank Reconstruction for Semantic Segmentation
- Unmixing based PAN guided fusion network for hyperspectral imagery
- Should You Go Deeper? Optimizing Convolutional Neural Network Architectures without Training by Receptive Field Analysis
- Rank Flow Embedding for Unsupervised and Semi-Supervised Manifold Learning
- BWCNN: Blink to Word, a Real-Time Convolutional Neural Network Approach
- LLIC: Large Receptive Field Transform Coding with Adaptive Weights for Learned Image Compression
- Double Refinement Network for Efficient Indoor Monocular Depth Estimation
- Parsing-based View-aware Embedding Network for Vehicle Re-Identification
- Periodic Residual Learning for Crowd Flow Forecasting
- Training speaker recognition systems with limited data
- NAS-VAD: Neural Architecture Search for Voice Activity Detection
- Depthwise Multiception Convolution for Reducing Network Parameters without Sacrificing Accuracy
- Real-time Fusion Network for RGB-D Semantic Segmentation Incorporating Unexpected Obstacle Detection for Road-driving Images
- Left Ventricle Quantification Using Direct Regression with Segmentation Regularization and Ensembles of Pretrained 2D and 3D CNNs
- Audio Anti-spoofing Using a Simple Attention Module and Joint Optimization Based on Additive Angular Margin Loss and Meta-learning
- Automated skin lesion segmentation using multi-scale feature extraction scheme and dual-attention mechanism
- Data Uncertainty Learning in Face Recognition
- CAGAN: Text-To-Image Generation with Combined Attention GANs
- A Comprehensive Study on Medical Image Segmentation using Deep Neural Networks
- Resolution-invariant Person Re-Identification
- Polynomial Neural Fields for Subband Decomposition and Manipulation
- Data and Knowledge Co-driving for Cancer Subtype Classification on Multi-Scale Histopathological Slides
- Headless Horseman: Adversarial Attacks on Transfer Learning Models
- Learning to Predict Context-adaptive Convolution for Semantic Segmentation
- Modeling plate and spring reverberation using a DSP-informed deep neural network
- 3D Teeth Reconstruction from Panoramic Radiographs using Neural Implicit Functions
- AMEIR: Automatic Behavior Modeling, Interaction Exploration and MLP Investigation in the Recommender System
- Hybrid CNN Based Attention with Category Prior for User Image Behavior Modeling
- SuperChat: Dialogue Generation by Transfer Learning from Vision to Language using Two-dimensional Word Embedding and Pretrained ImageNet CNN Models
- All Attention U-NET for Semantic Segmentation of Intracranial Hemorrhages In Head CT Images
- Global Information Guided Video Anomaly Detection
- Challenge report: Recognizing Families In the Wild Data Challenge
- Echofilter: A Deep Learning Segmentation Model Improves the Automation, Standardization, and Timeliness for Post-Processing Echosounder Data in Tidal Energy Streams
- GASP: Gated Attention For Saliency Prediction
- KinePose: A temporally optimized inverse kinematics technique for 6DOF human pose estimation with biomechanical constraints
- Accelerating temporal action proposal generation via high performance computing
- Learning to Segment Human Body Parts with Synthetically Trained Deep Convolutional Networks
- Bidirectional Knowledge Reconfiguration for Lightweight Point Cloud Analysis
- Spatial-aware Speaker Diarization for Multi-channel Multi-party Meeting
- CEKD:Cross Ensemble Knowledge Distillation for Augmented Fine-grained Data
- Evolving Neural Selection with Adaptive Regularization
- Contextual Classification Using Self-Supervised Auxiliary Models for Deep Neural Networks
- Trapped in texture bias? A large scale comparison of deep instance segmentation
- AIO-P: Expanding Neural Performance Predictors Beyond Image Classification
- End-to-End Cascaded U-Nets with a Localization Network for Kidney Tumor Segmentation
- Residual Channel Attention Network for Brain Glioma Segmentation
- WeightNet: Revisiting the Design Space of Weight Networks
- Bounding-box deep calibration for high performance face detection
- PhytNet -- Tailored Convolutional Neural Networks for Custom Botanical Data
- Balanced Binary Neural Networks with Gated Residual
- Temporal Self-Ensembling Teacher for Semi-Supervised Object Detection
- Deep Spectro-temporal Artifacts for Detecting Synthesized Speech
- Eye Semantic Segmentation with a Lightweight Model
- Deep High-Resolution Network for Low Dose X-ray CT Denoising
- TS-CAM: Token Semantic Coupled Attention Map for Weakly Supervised Object Localization
- Automatic rating of incomplete hippocampal inversions evaluated across multiple cohorts
- Weakly- and Semi-Supervised Probabilistic Segmentation and Quantification of Ultrasound Needle-Reverberation Artifacts to Allow Better AI Understanding of Tissue Beneath Needles
- Multi-QuartzNet: Multi-Resolution Convolution for Speech Recognition with Multi-Layer Feature Fusion
- End-to-end analysis using image classification
- SIPA: A Simple Framework for Efficient Networks
- Context Prior for Scene Segmentation
- Set-Based Face Recognition Beyond Disentanglement: Burstiness Suppression With Variance Vocabulary
- Multi-view Feature Augmentation with Adaptive Class Activation Mapping
- Efficient Modelling Across Time of Human Actions and Interactions
- Placepedia: Comprehensive Place Understanding with Multi-Faceted Annotations
- Net: Augmented Parallel-Pyramid Net for Attention Guided Pose Estimation