Multi-Scale Context Aggregation by Dilated Convolutions
arXiv:1511.07122
Abstract
State-of-the-art models for semantic segmentation are based on adaptations of convolutional networks that had originally been designed for image classification. However, dense prediction and image classification are structurally different. In this work, we develop a new convolutional network module that is specifically designed for dense prediction. The presented module uses dilated convolutions to systematically aggregate multi-scale contextual information without losing resolution. The architecture is based on the fact that dilated convolutions support exponential expansion of the receptive field without loss of resolution or coverage. We show that the presented context module increases the accuracy of state-of-the-art semantic segmentation systems. In addition, we examine the adaptation of image classification networks to dense prediction and show that simplifying the adapted network can increase accuracy.
Published as a conference paper at ICLR 2016
References in corpus (2)
Cited by in corpus (191)
- What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?
- Semantic Foggy Scene Understanding with Synthetic Data
- Fully Convolutional Networks for Semantic Segmentation
- Deep Bilateral Learning for Real-Time Image Enhancement
- BACH: Grand Challenge on Breast Cancer Histology Images
- FCNs in the Wild: Pixel-level Adversarial and Constraint-based Adaptation
- Automatic lung segmentation in routine imaging is primarily a data diversity problem, not a methodology problem
- Semantic Instance Segmentation with a Discriminative Loss Function
- Deformable Convolutional Networks
- Mapping the Landscape of Artificial Intelligence Applications against COVID-19
- MILD-Net: Minimal Information Loss Dilated Network for Gland Instance Segmentation in Colon Histology Images
- Pyramid Scene Parsing Network
- Semantic Labeling in Very High Resolution Images via a Self-Cascaded Convolutional Neural Network
- X-ModalNet: A Semi-Supervised Deep Cross-Modal Network for Classification of Remote Sensing Data
- Var-CNN: A Data-Efficient Website Fingerprinting Attack Based on Deep Learning
- A Survey on Deep Learning-based Architectures for Semantic Segmentation on 2D images
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- Self-Supervised Model Adaptation for Multimodal Semantic Segmentation
- Weakly Supervised Adversarial Domain Adaptation for Semantic Segmentation in Urban Scenes
- Regularization for Deep Learning: A Taxonomy
- Enhanced CNN for image denoising
- Why do deep convolutional networks generalize so poorly to small image transformations?
- Dilated Residual Networks
- Automatic Brain Tumor Segmentation using Convolutional Neural Networks with Test-Time Augmentation
- Graph Neural Networks: Taxonomy, Advances and Trends
- A Comprehensive Overview and Comparative Analysis on Deep Learning Models: CNN, RNN, LSTM, GRU
- An automatic COVID-19 CT segmentation network using spatial and channel attention mechanism
- Maximum Classifier Discrepancy for Unsupervised Domain Adaptation
- Deformable Kernel Networks for Joint Image Filtering
- Rain Removal in Traffic Surveillance: Does it Matter?
- Artistic style transfer for videos and spherical images
- Adapting Mask-RCNN for Automatic Nucleus Segmentation
- Implicit Dual-domain Convolutional Network for Robust Color Image Compression Artifact Reduction
- DeepGCNs: Can GCNs Go as Deep as CNNs?
- PointSeg: Real-Time Semantic Segmentation Based on 3D LiDAR Point Cloud
- DSOD: Learning Deeply Supervised Object Detectors from Scratch
- PixelNet: Representation of the pixels, by the pixels, and for the pixels
- Disentangling and Unifying Graph Convolutions for Skeleton-Based Action Recognition
- Learning Deep Context-aware Features over Body and Latent Parts for Person Re-identification
- Learning a Discriminative Feature Network for Semantic Segmentation
- SkinNet: A Deep Learning Framework for Skin Lesion Segmentation
- Photographic Image Synthesis with Cascaded Refinement Networks
- Gated-Dilated Networks for Lung Nodule Classification in CT scans
- Pyramid Stereo Matching Network
- RefineNet: Multi-Path Refinement Networks for High-Resolution Semantic Segmentation
- Learning to Denoise Astronomical Images with U-nets
- Learning Feature Pyramids for Human Pose Estimation
- Spatio-Temporal Graph Neural Point Process for Traffic Congestion Event Prediction
- CDC: Convolutional-De-Convolutional Networks for Precise Temporal Action Localization in Untrimmed Videos
- Improving Fast Segmentation With Teacher-student Learning
- ExFuse: Enhancing Feature Fusion for Semantic Segmentation
- Colorization as a Proxy Task for Visual Understanding
- Waterfall Atrous Spatial Pooling Architecture for Efficient Semantic Segmentation
- Distance transform regression for spatially-aware deep semantic segmentation
- Fast Image Processing with Fully-Convolutional Networks
- Fully Convolutional Networks for Chip-wise Defect Detection Employing Photoluminescence Images
- Deep Watershed Transform for Instance Segmentation
- Tracklets Predicting Based Adaptive Graph Tracking
- Rethinking the Faster R-CNN Architecture for Temporal Action Localization
- Doubly Convolutional Neural Networks
- DeepUNet: A Deep Fully Convolutional Network for Pixel-level Sea-Land Segmentation
- Simple Does It: Weakly Supervised Instance and Semantic Segmentation
- Convolutional Neural Pyramid for Image Processing
- Exploiting saliency for object segmentation from image level labels
- Semantic segmentation of mFISH images using convolutional networks
- Neural Nearest Neighbors Networks
- Multiple Instance Detection Network with Online Instance Classifier Refinement
- Video Propagation Networks
- Models Matter, So Does Training: An Empirical Study of CNNs for Optical Flow Estimation
- No More Discrimination: Cross City Adaptation of Road Scene Segmenters
- Image Inpainting using Block-wise Procedural Training with Annealed Adversarial Counterpart
- Electricity Theft Detection with self-attention
- Temporal Convolutional Networks for Action Segmentation and Detection
- Provably scale-covariant continuous hierarchical networks based on scale-normalized differential expressions coupled in cascade
- An End-to-End Network for Panoptic Segmentation
- Mutual Graph Learning for Camouflaged Object Detection
- A Novel Upsampling and Context Convolution for Image Semantic Segmentation
- Distribution-aware Margin Calibration for Semantic Segmentation in Images
- Agile Amulet: Real-Time Salient Object Detection with Contextual Attention
- Knowledge Adaptation for Efficient Semantic Segmentation
- Semantic Scene Completion Combining Colour and Depth: preliminary experiments
- All about Structure: Adapting Structural Information across Domains for Boosting Semantic Segmentation
- Cephalometric Landmark Detection by AttentiveFeature Pyramid Fusion and Regression-Voting
- Comparison of Deep learning models on time series forecasting : a case study of Dissolved Oxygen Prediction
- End-to-end semantic face segmentation with conditional random fields as convolutional, recurrent and adversarial networks
- Vision Transformers: From Semantic Segmentation to Dense Prediction
- Efficient Visual Recognition with Deep Neural Networks: A Survey on Recent Advances and New Directions
- LabelBank: Revisiting Global Perspectives for Semantic Segmentation
- Autofocus Layer for Semantic Segmentation
- Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes
- FETNet: Feature Erasing and Transferring Network for Scene Text Removal
- Hierarchical Discrete Distribution Decomposition for Match Density Estimation
- Learning to Synthesize a 4D RGBD Light Field from a Single Image
- Adaptive Semantic Segmentation with a Strategic Curriculum of Proxy Labels
- Scene Parsing with Global Context Embedding
- Object Detection Free Instance Segmentation With Labeling Transformations
- Star Shape Prior in Fully Convolutional Networks for Skin Lesion Segmentation
- Single Image Reflection Separation with Perceptual Losses
- Motion-Appearance Interactive Encoding for Object Segmentation in Unconstrained Videos
- A 2D dilated residual U-Net for multi-organ segmentation in thoracic CT
- Non-locally Enhanced Encoder-Decoder Network for Single Image De-raining
- ROI-Aware Multiscale Cross-Attention Vision Transformer for Pest Image Identification
- Learning Deep Representations for Semantic Image Parsing: a Comprehensive Overview
- WordFence: Text Detection in Natural Images with Border Awareness
- Actor-Action Semantic Segmentation with Region Masks
- Semantic Correlation Promoted Shape-Variant Context for Segmentation
- Semantically Consistent Image Completion with Fine-grained Details
- Learning a Dilated Residual Network for SAR Image Despeckling
- Pointwise Convolutional Neural Networks
- Compact Global Descriptor for Neural Networks
- Efficient Road Lane Marking Detection with Deep Learning
- Table Structure Recognition using Top-Down and Bottom-Up Cues
- DFPENet-geology: A Deep Learning Framework for High Precision Recognition and Segmentation of Co-seismic Landslides
- BowNet: Dilated Convolution Neural Network for Ultrasound Tongue Contour Extraction
- Automated Detecting and Placing Road Objects from Street-level Images
- Unsupervised Single Image Deraining with Self-supervised Constraints
- Super-Resolution with Deep Adaptive Image Resampling
- An Attention-Based System for Damage Assessment Using Satellite Imagery
- Deep Convolutional Neural Networks with Spatial Regularization, Volume and Star-shape Priori for Image Segmentation
- KeyPose: Multi-View 3D Labeling and Keypoint Estimation for Transparent Objects
- Data Stream Stabilization for Optical Coherence Tomography Volumetric Scanning
- Adaptive Weighting Multi-Field-of-View CNN for Semantic Segmentation in Pathology
- Dense Recurrent Neural Networks for Scene Labeling
- Learning to detect and localize many objects from few examples
- A Multiscale Patch Based Convolutional Network for Brain Tumor Segmentation
- Towards large-scale, automated, accurate detection of CCTV camera objects using computer vision. Applications and implications for privacy, safety, and cybersecurity. (Preprint)
- Artificial intelligence based prediction on lung cancer risk factors using deep learning
- D2A U-Net: Automatic Segmentation of COVID-19 Lesions from CT Slices with Dilated Convolution and Dual Attention Mechanism
- Weakly supervised training of universal visual concepts for multi-domain semantic segmentation
- Distilling Ensemble of Explanations for Weakly-Supervised Pre-Training of Image Segmentation Models
- Locally Adaptive Learning Loss for Semantic Image Segmentation
- A framework for the fine-grained evaluation of the instantaneous expected value of soccer possessions
- Scene Understanding Networks for Autonomous Driving based on Around View Monitoring System
- Accurate Automatic Segmentation of Amygdala Subnuclei and Modeling of Uncertainty via Bayesian Fully Convolutional Neural Network
- Teaching Machines to Code: Neural Markup Generation with Visual Attention
- Deep Multicameral Decoding for Localizing Unoccluded Object Instances from a Single RGB Image
- A Unified Efficient Pyramid Transformer for Semantic Segmentation
- Dual Encoder Fusion U-Net (DEFU-Net) for Cross-manufacturer Chest X-ray Segmentation
- Attention Mechanisms for Object Recognition with Event-Based Cameras
- Multiview Two-Task Recursive Attention Model for Left Atrium and Atrial Scars Segmentation
- Design and Development of Autonomous Delivery Robot
- Deep Learning for Low-Dose CT Denoising
- Cross-Domain Self-supervised Multi-task Feature Learning using Synthetic Imagery
- Identifying Most Walkable Direction for Navigation in an Outdoor Environment
- Physically-Based Rendering for Indoor Scene Understanding Using Convolutional Neural Networks
- What and Where: A Context-based Recommendation System for Object Insertion
- Improving the Resolution of CNN Feature Maps Efficiently with Multisampling
- AutoScaler: Scale-Attention Networks for Visual Correspondence
- ExplainFix: Explainable Spatially Fixed Deep Networks
- Reconstruction of Simulation-Based Physical Field by Reconstruction Neural Network Method
- Deep Dual Pyramid Network for Barcode Segmentation using Barcode-30k Database
- Fast ES-RNN: A GPU Implementation of the ES-RNN Algorithm
- Semantic-Unit-Based Dilated Convolution for Multi-Label Text Classification
- Automatic Brain Structures Segmentation Using Deep Residual Dilated U-Net
- Semantic Segmentation of Earth Observation Data Using Multimodal and Multi-scale Deep Networks
- Learning to Sieve: Prediction of Grading Curves from Images of Concrete Aggregate
- Object-oriented Neural Programming (OONP) for Document Understanding
- Spatially-Adaptive Filter Units for Deep Neural Networks
- Dissociating model architectures from inference computations
- Dynamic Video Segmentation Network
- Attentive CT Lesion Detection Using Deep Pyramid Inference with Multi-Scale Booster
- A Separable Temporal Convolution Neural Network with Attention for Small-Footprint Keyword Spotting
- Detection of cellular micromotion by advanced signal processing
- Backdrop: Stochastic Backpropagation
- Exposing Semantic Segmentation Failures via Maximum Discrepancy Competition
- A Novel Multi-scale Dilated 3D CNN for Epileptic Seizure Prediction
- U-Net with spatial pyramid pooling for drusen segmentation in optical coherence tomography
- A Domain Agnostic Normalization Layer for Unsupervised Adversarial Domain Adaptation
- Deep Multiple Description Coding by Learning Scalar Quantization
- Scene Parsing via Dense Recurrent Neural Networks with Attentional Selection
- Dilated Spatial Generative Adversarial Networks for Ergodic Image Generation
- Dense-Resolution Network for Point Cloud Classification and Segmentation
- Brain Graph Super-Resolution Using Adversarial Graph Neural Network with Application to Functional Brain Connectivity
- Deep Optimized Multiple Description Image Coding via Scalar Quantization Learning
- Compressed Sensing MRI via a Multi-scale Dilated Residual Convolution Network
- Dense Fusion Classmate Network for Land Cover Classification
- Automatic Liver Segmentation with Adversarial Loss and Convolutional Neural Network
- Encoder-Decoder based CNN and Fully Connected CRFs for Remote Sensed Image Segmentation
- Split-Merge Pooling
- Attentive Sequence to Sequence Translation for Localizing Clips of Interest by Natural Language Descriptions
- Dilated filters for edge detection algorithms
- Computationally efficient cardiac views projection using 3D Convolutional Neural Networks
- Knowledge-based Fully Convolutional Network and Its Application in Segmentation of Lung CT Images
- ReflectNet -- A Generative Adversarial Method for Single Image Reflection Suppression
- Soccer on Your Tabletop
- Learning Equivariant Representations
- Multi-Level Fine-Tuning: Closing Generalization Gaps in Approximation of Solution Maps under a Limited Budget for Training Data
- Semantic Segmentation for Urban-Scene Images
- Cyclic orthogonal convolutions for long-range integration of features
- Atlas-aware ConvNetfor Accurate yet Robust Anatomical Segmentation
- Synthesizing Photorealistic Images with Deep Generative Learning