DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
arXiv:1606.00915
Abstract
In this work we address the task of semantic image segmentation with Deep Learning and make three main contributions that are experimentally shown to have substantial practical merit. First, we highlight convolution with upsampled filters, or 'atrous convolution', as a powerful tool in dense prediction tasks. Atrous convolution allows us to explicitly control the resolution at which feature responses are computed within Deep Convolutional Neural Networks. It also allows us to effectively enlarge the field of view of filters to incorporate larger context without increasing the number of parameters or the amount of computation. Second, we propose atrous spatial pyramid pooling (ASPP) to robustly segment objects at multiple scales. ASPP probes an incoming convolutional feature layer with filters at multiple sampling rates and effective fields-of-views, thus capturing objects as well as image context at multiple scales. Third, we improve the localization of object boundaries by combining methods from DCNNs and probabilistic graphical models. The commonly deployed combination of max-pooling and downsampling in DCNNs achieves invariance but has a toll on localization accuracy. We overcome this by combining the responses at the final DCNN layer with a fully connected Conditional Random Field (CRF), which is shown both qualitatively and quantitatively to improve localization performance. Our proposed "DeepLab" system sets the new state-of-art at the PASCAL VOC-2012 semantic image segmentation task, reaching 79.7% mIOU in the test set, and advances the results on three other datasets: PASCAL-Context, PASCAL-Person-Part, and Cityscapes. All of our code is made publicly available online.
Accepted by TPAMI
References in corpus (17)
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Learning Deconvolution Network for Semantic Segmentation
- Fully Connected Deep Structured Networks
- Bridging Category-level and Instance-level Semantic Image Segmentation
- BoxSup: Exploiting Bounding Boxes to Supervise Convolutional Networks for Semantic Segmentation
- High-performance Semantic Segmentation Using Very Deep Fully Convolutional Networks
- Instance-sensitive Fully Convolutional Networks
- Pixel-level Encoding and Depth Layering for Instance-level Semantic Labeling
- Laplacian Pyramid Reconstruction and Refinement for Semantic Segmentation
- Joint Object and Part Segmentation using Deep Learned Potentials
- Material Recognition in the Wild with the Materials in Context Database
- Optical Flow with Semantic Segmentation and Localized Layers
- Combining the Best of Graphical Models and ConvNets for Semantic Segmentation
- From Image-level to Pixel-level Labeling with Convolutional Networks
- Fast Semantic Image Segmentation with High Order Context and Guided Filtering
- Semantic Object Parsing with Graph LSTM
Cited by in corpus (202)
- MobileNetV2: Inverted Residuals and Linear Bottlenecks
- ResUNet-a: a deep learning framework for semantic segmentation of remotely sensed data
- Momentum Contrast for Unsupervised Visual Representation Learning
- SegFormer: Simple and Efficient Design for Semantic Segmentation with Transformers
- A deep learning model integrating FCNNs and CRFs for brain tumor segmentation
- Learning Spectral-Spatial-Temporal Features via a Recurrent Convolutional Neural Network for Change Detection in Multispectral Imagery
- Rethinking Pre-training and Self-training
- Adversarial Learning for Semi-Supervised Semantic Segmentation
- On the Compactness, Efficiency, and Representation of 3D Convolutional Networks: Brain Parcellation as a Pretext Task
- Evolution of Image Segmentation using Deep Convolutional Neural Network: A Survey
- Pyramid Scene Parsing Network
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- FastFCN: Rethinking Dilated Convolution in the Backbone for Semantic Segmentation
- Convolutional neural networks automate detection for tracking of submicron scale particles in 2D and 3D
- Semantic Understanding of Scenes through the ADE20K Dataset
- Dilated Residual Networks
- Robust Spatial Filtering with Graph Convolutional Neural Networks
- Auto-DeepLab: Hierarchical Neural Architecture Search for Semantic Image Segmentation
- Virtual to Real Reinforcement Learning for Autonomous Driving
- Examining the Impact of Blur on Recognition by Convolutional Networks
- Context Encoding for Semantic Segmentation
- ICNet for Real-Time Semantic Segmentation on High-Resolution Images
- Image Segmentation Algorithms Overview
- BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation
- Automated sub-cortical brain structure segmentation combining spatial and deep convolutional features
- Light-Weight RefineNet for Real-Time Semantic Segmentation
- B-CNN: Branch Convolutional Neural Network for Hierarchical Classification
- Towards Image Understanding from Deep Compression without Decoding
- What Do We Understand About Convolutional Networks?
- PixelNet: Representation of the pixels, by the pixels, and for the pixels
- Adversarial Examples for Semantic Segmentation and Object Detection
- Learning a Discriminative Feature Network for Semantic Segmentation
- Fast Object Learning and Dual-arm Coordination for Cluttered Stowing, Picking, and Packing
- Towards Accurate Multi-person Pose Estimation in the Wild
- Gated-Dilated Networks for Lung Nodule Classification in CT scans
- Learning Video Object Segmentation from Static Images
- The One Hundred Layers Tiramisu: Fully Convolutional DenseNets for Semantic Segmentation
- Knowledge Distillation in Generations: More Tolerant Teachers Educate Better Students
- RefineNet: Multi-Path Refinement Networks for High-Resolution Semantic Segmentation
- CDC: Convolutional-De-Convolutional Networks for Precise Temporal Action Localization in Untrimmed Videos
- PixelNet: Towards a General Pixel-level Architecture
- Driving Scene Perception Network: Real-time Joint Detection, Depth Estimation and Semantic Segmentation
- Real-time Semantic Image Segmentation via Spatial Sparsity
- Train Large, Then Compress: Rethinking Model Size for Efficient Training and Inference of Transformers
- A Survey on Deep Learning Methods for Robot Vision
- Predicting Scene Parsing and Motion Dynamics in the Future
- Looking at Outfit to Parse Clothing
- Semi-Dense 3D Semantic Mapping from Monocular SLAM
- CUTIE: Learning to Understand Documents with Convolutional Universal Text Information Extractor
- Visual Affordance and Function Understanding: A Survey
- SRM : A Style-based Recalibration Module for Convolutional Neural Networks
- Learning to Branch for Multi-Task Learning
- The Lovász-Softmax loss: A tractable surrogate for the optimization of the intersection-over-union measure in neural networks
- Domain Adaptation for Semantic Segmentation with Maximum Squares Loss
- ResizeMix: Mixing Data with Preserved Object Information and True Labels
- DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting
- Checkmate: Breaking the Memory Wall with Optimal Tensor Rematerialization
- MaskLab: Instance Segmentation by Refining Object Detection with Semantic and Direction Features
- Interactive Video Object Segmentation in the Wild
- Scene Text Synthesis for Efficient and Effective Deep Network Training
- All You Need is Beyond a Good Init: Exploring Better Solution for Training Extremely Deep Convolutional Neural Networks with Orthonormality and Modulation
- Convolutional Neural Pyramid for Image Processing
- Exploiting saliency for object segmentation from image level labels
- Pixel Deconvolutional Networks
- Getting to 99% Accuracy in Interactive Segmentation
- FEELVOS: Fast End-to-End Embedding Learning for Video Object Segmentation
- Learning Affinity via Spatial Propagation Networks
- Learning Implicitly Recurrent CNNs Through Parameter Sharing
- Not All Pixels Are Equal: Difficulty-aware Semantic Segmentation via Deep Layer Cascade
- Deep Learning for Pancreas Segmentation: a Systematic Review
- Multi-Scale Coarse-to-Fine Segmentation for Screening Pancreatic Ductal Adenocarcinoma
- Macroscale fracture surface segmentation via semi-supervised learning considering the structural similarity
- DCNAS: Densely Connected Neural Architecture Search for Semantic Image Segmentation
- Fashioning with Networks: Neural Style Transfer to Design Clothes
- On Regularized Losses for Weakly-supervised CNN Segmentation
- CurriculumNet: Weakly Supervised Learning from Large-Scale Web Images
- Convolutional Video Steganography with Temporal Residual Modeling
- All about Structure: Adapting Structural Information across Domains for Boosting Semantic Segmentation
- Exploit fully automatic low-level segmented PET data for training high-level deep learning algorithms for the corresponding CT data
- Counterfactual Critic Multi-Agent Training for Scene Graph Generation
- Cost Volume Pyramid Based Depth Inference for Multi-View Stereo
- Efficient Accelerator for Dilated and Transposed Convolution with Decomposition
- Mixup-CAM: Weakly-supervised Semantic Segmentation via Uncertainty Regularization
- Blazingly Fast Video Object Segmentation with Pixel-Wise Metric Learning
- Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes
- CascadePSP: Toward Class-Agnostic and Very High-Resolution Segmentation via Global and Local Refinement
- Correlation Maximized Structural Similarity Loss for Semantic Segmentation
- Object Detection Free Instance Segmentation With Labeling Transformations
- Joint Graph Decomposition and Node Labeling: Problem, Algorithms, Applications
- Scene Parsing with Global Context Embedding
- A Relation-Augmented Fully Convolutional Network for Semantic Segmentation in Aerial Scenes
- Simpler is Better: Few-shot Semantic Segmentation with Classifier Weight Transformer
- Deep Extreme Cut: From Extreme Points to Object Segmentation
- Star Shape Prior in Fully Convolutional Networks for Skin Lesion Segmentation
- Cross-domain Human Parsing via Adversarial Feature and Label Adaptation
- Learning Dynamic Routing for Semantic Segmentation
- Learning Deep Representations for Semantic Image Parsing: a Comprehensive Overview
- Learning Rigidity in Dynamic Scenes with a Moving Camera for 3D Motion Field Estimation
- An Embarrassingly Simple Approach for Knowledge Distillation
- Actor-Action Semantic Segmentation with Region Masks
- Pointwise Convolutional Neural Networks
- Weakly Supervised Semantic Segmentation by Pixel-to-Prototype Contrast
- Weakly-Supervised Semantic Segmentation via Sub-category Exploration
- CO2: Consistent Contrast for Unsupervised Visual Representation Learning
- FoveaNet: Perspective-aware Urban Scene Parsing
- Weakly Supervised Semantic Segmentation Based on Web Image Co-segmentation
- MSG-Transformer: Exchanging Local Spatial Information by Manipulating Messenger Tokens
- LiftPool: Bidirectional ConvNet Pooling
- Zero-Shot Semantic Segmentation
- AdaCoSeg: Adaptive Shape Co-Segmentation with Group Consistency Loss
- Relating Input Concepts to Convolutional Neural Network Decisions
- Look Before You Leap: Learning Landmark Features for One-Stage Visual Grounding
- STEP: Segmenting and Tracking Every Pixel
- Accelerating Deep Neural Networks with Spatial Bottleneck Modules
- Super-Resolution with Deep Adaptive Image Resampling
- Referring Expression Object Segmentation with Caption-Aware Consistency
- Explorations and Lessons Learned in Building an Autonomous Formula SAE Car from Simulations
- A 3D Coarse-to-Fine Framework for Volumetric Medical Image Segmentation
- Bi-Mix: Bidirectional Mixing for Domain Adaptive Nighttime Semantic Segmentation
- Dense Recurrent Neural Networks for Scene Labeling
- Weakly and Semi Supervised Human Body Part Parsing via Pose-Guided Knowledge Transfer
- A deep learning-based method for prostate segmentation in T2-weighted magnetic resonance imaging
- SketchyScene: Richly-Annotated Scene Sketches
- Boundary-sensitive Network for Portrait Segmentation
- Adaptive Binarization for Weakly Supervised Affordance Segmentation
- Sparsely Aggregated Convolutional Networks
- Beyond Forward Shortcuts: Fully Convolutional Master-Slave Networks (MSNets) with Backward Skip Connections for Semantic Segmentation
- Learning Multi-modal Information for Robust Light Field Depth Estimation
- RefineMask: Towards High-Quality Instance Segmentation with Fine-Grained Features
- The Resistance to Label Noise in K-NN and DNN Depends on its Concentration
- Multigrid Neural Architectures
- Sketch2code: Generating a website from a paper mockup
- LUCSS: Language-based User-customized Colourization of Scene Sketches
- BusyHands: A Hand-Tool Interaction Database for Assembly Tasks Semantic Segmentation
- Unsupervised domain adaptation via coarse-to-fine feature alignment method using contrastive learning
- Dual Encoder Fusion U-Net (DEFU-Net) for Cross-manufacturer Chest X-ray Segmentation
- Unsupervised Domain Adaptation for Video Semantic Segmentation
- Stacked Neural Networks for end-to-end ciliary motion analysis
- Outline Objects using Deep Reinforcement Learning
- SeGAN: Segmenting and Generating the Invisible
- A Unified Efficient Pyramid Transformer for Semantic Segmentation
- EfficientHRNet: Efficient Scaling for Lightweight High-Resolution Multi-Person Pose Estimation
- Gotta Adapt 'Em All: Joint Pixel and Feature-Level Domain Adaptation for Recognition in the Wild
- Referring Image Segmentation via Cross-Modal Progressive Comprehension
- AutoLoss-Zero: Searching Loss Functions from Scratch for Generic Tasks
- MLPerf HPC: A Holistic Benchmark Suite for Scientific Machine Learning on HPC Systems
- Ray-ONet: Efficient 3D Reconstruction From A Single RGB Image
- Learning Multiple Dense Prediction Tasks from Partially Annotated Data
- Image Labeling with Markov Random Fields and Conditional Random Fields
- Influence of Image Classification Accuracy on Saliency Map Estimation
- Improving the Resolution of CNN Feature Maps Efficiently with Multisampling
- Binge Watching: Scaling Affordance Learning from Sitcoms
- FaceShapeGene: A Disentangled Shape Representation for Flexible Face Image Editing
- Face Aging with Contextual Generative Adversarial Nets
- Spatial Memory for Context Reasoning in Object Detection
- Rethinking Convolutional Semantic Segmentation Learning
- Inter-BMV: Interpolation with Block Motion Vectors for Fast Semantic Segmentation on Video
- Patchwork: A Patch-wise Attention Network for Efficient Object Detection and Segmentation in Video Streams
- Spatially-Adaptive Filter Units for Deep Neural Networks
- Quality-Aware Network for Human Parsing
- Parsing R-CNN for Instance-Level Human Analysis
- IMP: Instance Mask Projection for High Accuracy Semantic Segmentation of Things
- Representative Graph Neural Network
- Improving Augmentation and Evaluation Schemes for Semantic Image Synthesis
- Scene Parsing via Dense Recurrent Neural Networks with Attentional Selection
- Dilated Spatial Generative Adversarial Networks for Ergodic Image Generation
- The Ethical Dilemma when (not) Setting up Cost-based Decision Rules in Semantic Segmentation
- Diagnostics in Semantic Segmentation
- S4-Net: Geometry-Consistent Semi-Supervised Semantic Segmentation
- Unseen Object Segmentation in Videos via Transferable Representations
- Structured 2D Representation of 3D Data for Shape Processing
- Unsupervised Video Object Segmentation with Distractor-Aware Online Adaptation
- Context Prior for Scene Segmentation
- Detection and Classification of Breast Cancer Metastates Based on U-Net
- Design Pseudo Ground Truth with Motion Cue for Unsupervised Video Object Segmentation
- Beyond Planar Symmetry: Modeling human perception of reflection and rotation symmetries in the wild
- PointINS: Point-based Instance Segmentation
- What Can You Learn from Your Muscles? Learning Visual Representation from Human Interactions
- Recurrent Multimodal Interaction for Referring Image Segmentation
- AutoDrop: Training Deep Learning Models with Automatic Learning Rate Drop
- Image Synthesis via Semantic Composition
- Universal Perceptual Grouping
- Deep, Dense, and Low-Rank Gaussian Conditional Random Fields
- Shelf-Supervised Mesh Prediction in the Wild
- A Survey On 3D Inner Structure Prediction from its Outer Shape
- WeedMap: A large-scale semantic weed mapping framework using aerial multispectral imaging and deep neural network for precision farming
- Multi-Scale Spatially-Asymmetric Recalibration for Image Classification
- Semantic See-Through Rendering on Light Fields
- Guided Feature Selection for Deep Visual Odometry
- Addressing the Invisible: Street Address Generation for Developing Countries with Deep Learning
- 3D Scene Parsing via Class-Wise Adaptation
- RethNet: Object-by-Object Learning for Detecting Facial Skin Problems
- Beyond Gradient Descent for Regularized Segmentation Losses
- CNN in MRF: Video Object Segmentation via Inference in A CNN-Based Higher-Order Spatio-Temporal MRF
- MASON: A Model AgnoStic ObjectNess Framework
- Using Cross-Model EgoSupervision to Learn Cooperative Basketball Intention
- Per-Pixel Feedback for improving Semantic Segmentation
- A Distraction Score for Watermarks
- A Study on Trees's Knots Prediction from their Bark Outer-Shape
- DV3+HED+: A DCNNs-based Framework to Monitor Temporary Works and ESAs in Railway Construction Project Using VHR Satellite Images
- Learning Rich Representations For Structured Visual Prediction Tasks
- Quality-Aware Network for Face Parsing