Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs
arXiv:1412.7062
Abstract
Deep Convolutional Neural Networks (DCNNs) have recently shown state of the art performance in high level vision tasks, such as image classification and object detection. This work brings together methods from DCNNs and probabilistic graphical models for addressing the task of pixel-level classification (also called "semantic image segmentation"). We show that responses at the final layer of DCNNs are not sufficiently localized for accurate object segmentation. This is due to the very invariance properties that make DCNNs good for high level tasks. We overcome this poor localization property of deep networks by combining the responses at the final DCNN layer with a fully connected Conditional Random Field (CRF). Qualitatively, our "DeepLab" system is able to localize segment boundaries at a level of accuracy which is beyond previous methods. Quantitatively, our method sets the new state-of-art at the PASCAL VOC-2012 semantic image segmentation task, reaching 71.6% IOU accuracy in the test set. We show how these results can be obtained efficiently: Careful network re-purposing and a novel application of the 'hole' algorithm from the wavelet community allow dense computation of neural net responses at 8 frames per second on a modern GPU.
14 pages. Updated related work
References in corpus (18)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Fully Convolutional Networks for Semantic Segmentation
- Conditional Random Fields as Recurrent Neural Networks
- Going Deeper with Convolutions
- Joint Training of a Convolutional Network and a Graphical Model for Human Pose Estimation
- Fully Convolutional Networks for Semantic Segmentation
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- Convolutional Feature Masking for Joint Object and Stuff Segmentation
- Weakly- and Semi-Supervised Learning of a DCNN for Semantic Image Segmentation
- Articulated Pose Estimation by a Graphical Model with Image Dependent Pairwise Relations
- Simultaneous Detection and Segmentation
- Learning Deep Structured Models
- Feedforward semantic segmentation with zoom-out features
- Untangling Local and Global Deformations in Deep Convolutional Networks for Image Classification and Sliding Window Detection
- Material Recognition in the Wild with the Materials in Context Database
- Combining the Best of Graphical Models and ConvNets for Semantic Segmentation
Cited by in corpus (258)
- Efficient Multi-Scale 3D CNN with Fully Connected CRF for Accurate Brain Lesion Segmentation
- Conditional Random Fields as Recurrent Neural Networks
- Multi-Scale Context Aggregation by Dilated Convolutions
- Deep Multi-modal Object Detection and Semantic Segmentation for Autonomous Driving: Datasets, Methods, and Challenges
- ParseNet: Looking Wider to See Better
- Interactive Medical Image Segmentation using Deep Learning with Image-specific Fine-tuning
- Fully Convolutional Networks for Semantic Segmentation
- A deep learning model integrating FCNNs and CRFs for brain tumor segmentation
- Richer Convolutional Features for Edge Detection
- FCNs in the Wild: Pixel-level Adversarial and Constraint-based Adaptation
- Automatic Liver and Lesion Segmentation in CT Using Cascaded Fully Convolutional Neural Networks and 3D Conditional Random Fields
- SegAN: Adversarial Network with Multi-scale Loss for Medical Image Segmentation
- STC: A Simple to Complex Framework for Weakly-supervised Semantic Segmentation
- Learning Deconvolution Network for Semantic Segmentation
- DeepSaliency: Multi-Task Deep Neural Network Model for Salient Object Detection
- Curriculum Domain Adaptation for Semantic Segmentation of Urban Scenes
- Deep Label Distribution Learning with Label Ambiguity
- DeepIGeoS: A Deep Interactive Geodesic Framework for Medical Image Segmentation
- Classifying and Segmenting Microscopy Images Using Convolutional Multiple Instance Learning
- ABCNet: Attentive Bilateral Contextual Network for Efficient Semantic Segmentation of Fine-Resolution Remote Sensing Images
- Deformable Convolutional Networks
- 3D fully convolutional networks for subcortical segmentation in MRI: A large-scale study
- Fast-SCNN: Fast Semantic Segmentation Network
- DeepNAT: Deep Convolutional Neural Network for Segmenting Neuroanatomy
- Weakly- and Semi-Supervised Learning of a DCNN for Semantic Image Segmentation
- Evolution of Image Segmentation using Deep Convolutional Neural Network: A Survey
- Neural Machine Translation in Linear Time
- Semantic Labeling in Very High Resolution Images via a Self-Cascaded Convolutional Neural Network
- Multi-stage Attention ResU-Net for Semantic Segmentation of Fine-Resolution Remote Sensing Images
- Constrained-CNN losses for weakly supervised segmentation
- RANet: Ranking Attention Network for Fast Video Object Segmentation
- Segmentation of Glioma Tumors in Brain Using Deep Convolutional Neural Network
- Deep Dual-resolution Networks for Real-time and Accurate Semantic Segmentation of Road Scenes
- Modality specific U-Net variants for biomedical image segmentation: A survey
- Radar-Camera Fusion for Object Detection and Semantic Segmentation in Autonomous Driving: A Comprehensive Review
- A Survey on Deep Learning-based Architectures for Semantic Segmentation on 2D images
- Deep Learning for Generic Object Detection: A Survey
- Global Guidance Network for Breast Lesion Segmentation in Ultrasound Images
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- GridMask Data Augmentation
- Strengths and Weaknesses of Deep Learning Models for Face Recognition Against Image Degradations
- CrackGAN: Pavement Crack Detection Using Partially Accurate Ground Truths Based on Generative Adversarial Learning
- Progressive LiDAR Adaptation for Road Detection
- GhostNets on Heterogeneous Devices via Cheap Operations
- Automatic Liver and Tumor Segmentation of CT and MRI Volumes using Cascaded Fully Convolutional Neural Networks
- SNE-RoadSeg: Incorporating Surface Normal Information into Semantic Segmentation for Accurate Freespace Detection
- Semantic Image Segmentation via Deep Parsing Network
- Hybrid CNN and Dictionary-Based Models for Scene Recognition and Domain Adaptation
- RGB-D Object Detection and Semantic Segmentation for Autonomous Manipulation in Clutter
- Affinity Attention Graph Neural Network for Weakly Supervised Semantic Segmentation
- FusionSeg: Learning to combine motion and appearance for fully automatic segmention of generic objects in videos
- Non-local Neural Networks
- X-Net: Brain Stroke Lesion Segmentation Based on Depthwise Separable Convolution and Long-range Dependencies
- Spinal cord gray matter segmentation using deep dilated convolutions
- Context Encoding for Semantic Segmentation
- Shallow and Deep Convolutional Networks for Saliency Prediction
- DCAN: Deep Contour-Aware Networks for Accurate Gland Segmentation
- Edge Preserving and Multi-Scale Contextual Neural Network for Salient Object Detection
- Context-Aware Interaction Network for RGB-T Semantic Segmentation
- BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation
- DDU-Net: Dual-Decoder-U-Net for Road Extraction Using High-Resolution Remote Sensing Images
- Face Detection through Scale-Friendly Deep Convolutional Networks
- Single-Shot Refinement Neural Network for Object Detection
- Light-Weight RefineNet for Real-Time Semantic Segmentation
- Training Convolutional Neural Networks with Limited Training Data for Ear Recognition in the Wild
- Flow-Guided Feature Aggregation for Video Object Detection
- A Survey of Knowledge Representation in Service Robotics
- TransCAM: Transformer Attention-based CAM Refinement for Weakly Supervised Semantic Segmentation
- DSOD: Learning Deeply Supervised Object Detectors from Scratch
- Face Clustering: Representation and Pairwise Constraints
- Learning a Discriminative Feature Network for Semantic Segmentation
- SAC-Net: Spatial Attenuation Context for Salient Object Detection
- Learning Collision-Free Space Detection from Stereo Images: Homography Matrix Brings Better Data Augmentation
- T-Net: Nested encoder-decoder architecture for the main vessel segmentation in coronary angiography
- Segmentation in large-scale cellular electron microscopy with deep learning: A literature survey
- Detection and Classification of Supernova Gravitational Waves Signals: A Deep Learning Approach
- Recurrent Neural Networks to Correct Satellite Image Classification Maps
- Cross-Domain Complementary Learning Using Pose for Multi-Person Part Segmentation
- Subsurface structure analysis using computational interpretation and learning: A visual signal processing perspective
- Depth Adaptive Deep Neural Network for Semantic Segmentation
- MudrockNet: Semantic Segmentation of Mudrock SEM Images through Deep Learning
- GeoNet++: Iterative Geometric Neural Network with Edge-Aware Refinement for Joint Depth and Surface Normal Estimation
- Stage-Aware Feature Alignment Network for Real-Time Semantic Segmentation of Street Scenes
- Pipe-SGD: A Decentralized Pipelined SGD Framework for Distributed Deep Net Training
- A Radar Signal Deinterleaving Method Based on Semantic Segmentation with Neural Network
- PAD-Net: Multi-Tasks Guided Prediction-and-Distillation Network for Simultaneous Depth Estimation and Scene Parsing
- Learning like a Child: Fast Novel Visual Concept Learning from Sentence Descriptions of Images
- Toward quantitative fractography using convolutional neural networks
- Multi-Scale Feature Fusion: Learning Better Semantic Segmentation for Road Pothole Detection
- Learning to Refine Object Segments
- CDC: Convolutional-De-Convolutional Networks for Precise Temporal Action Localization in Untrimmed Videos
- Poverty Prediction with Public Landsat 7 Satellite Imagery and Machine Learning
- Video Semantic Segmentation with Distortion-Aware Feature Correction
- Development of Skip Connection in Deep Neural Networks for Computer Vision and Medical Image Analysis: A Survey
- Efficient Multi-Task RGB-D Scene Analysis for Indoor Environments
- Training Deep Networks with Structured Layers by Matrix Backpropagation
- Learning Video Object Segmentation with Visual Memory
- A Transformer-based Generative Adversarial Network for Brain Tumor Segmentation
- Real-Time Semantic Segmentation via Multiply Spatial Fusion Network
- Deep Weakly-Supervised Learning Methods for Classification and Localization in Histology Images: A Survey
- Deep Multi-Branch Aggregation Network for Real-Time Semantic Segmentation in Street Scenes
- The Stixel world: A medium-level representation of traffic scenes
- Adversarial Deep Structural Networks for Mammographic Mass Segmentation
- Attention guided global enhancement and local refinement network for semantic segmentation
- Incorporating Network Built-in Priors in Weakly-supervised Semantic Segmentation
- Annotating Object Instances with a Polygon-RNN
- Multi-Modal Obstacle Detection in Unstructured Environments with Conditional Random Fields
- Using Deep Learning to Localize Gravitational Wave Sources
- Structured Attention Guided Convolutional Neural Fields for Monocular Depth Estimation
- Guided Upsampling Network for Real-Time Semantic Segmentation
- Two-Phase Learning for Weakly Supervised Object Localization
- Learning from Pixel-Level Label Noise: A New Perspective for Semi-Supervised Semantic Segmentation
- Deep Semantic Segmentation at the Edge for Autonomous Navigation in Vineyard Rows
- Semantic Image Segmentation with Task-Specific Edge Detection Using CNNs and a Discriminatively Trained Domain Transform
- Weakly-Supervised Semantic Segmentation by Iteratively Mining Common Object Features
- A Learning Framework for n-bit Quantized Neural Networks toward FPGAs
- Predicting Deeper into the Future of Semantic Segmentation
- DeepCut: Joint Subset Partition and Labeling for Multi Person Pose Estimation
- SFD: Single Shot Scale-invariant Face Detector
- WarpNet: Weakly Supervised Matching for Single-view Reconstruction
- Stacked U-Nets: A No-Frills Approach to Natural Image Segmentation
- PDNet: Semantic Segmentation integrated with a Primal-Dual Network for Document binarization
- Simple Does It: Weakly Supervised Instance and Semantic Segmentation
- A Self-Distillation Embedded Supervised Affinity Attention Model for Few-Shot Segmentation
- BriNet: Towards Bridging the Intra-class and Inter-class Gaps in One-Shot Segmentation
- A Deep One-Shot Network for Query-based Logo Retrieval
- Built-in Foreground/Background Prior for Weakly-Supervised Semantic Segmentation
- Joint Object and Part Segmentation using Deep Learned Potentials
- Video Propagation Networks
- Real-Time and Accurate Object Detection in Compressed Video by Long Short-term Feature Aggregation
- Local Label Point Correction for Edge Detection of Overlapping Cervical Cells
- Segmenting Transparent Object in the Wild with Transformer
- CASENet: Deep Category-Aware Semantic Edge Detection
- Discriminative Training of Deep Fully-connected Continuous CRF with Task-specific Loss
- Evaluating Contrastive Models for Instance-based Image Retrieval
- CSC-Unet: A Novel Convolutional Sparse Coding Strategy Based Neural Network for Semantic Segmentation
- RSI-Net: Two-Stream Deep Neural Network for Remote Sensing Imagesbased Semantic Segmentation
- No More Discrimination: Cross City Adaptation of Road Scene Segmenters
- Improving Fully Convolution Network for Semantic Segmentation
- ORDNet: Capturing Omni-Range Dependencies for Scene Parsing
- A Novel Upsampling and Context Convolution for Image Semantic Segmentation
- Real-Time Facial Segmentation and Performance Capture from RGB Input
- Efficient Yet Deep Convolutional Neural Networks for Semantic Segmentation
- Optical Flow Estimation using a Spatial Pyramid Network
- Instance-Level Segmentation for Autonomous Driving with Deep Densely Connected MRFs
- Global Aggregation then Local Distribution for Scene Parsing
- Deep Image Harmonization
- Learning to Segment Instances in Videos with Spatial Propagation Network
- Label-guided Attention Distillation for Lane Segmentation
- Vision Transformers: From Semantic Segmentation to Dense Prediction
- LOANet: A Lightweight Network Using Object Attention for Extracting Buildings and Roads from UAV Aerial Remote Sensing Images
- Auto-context Convolutional Neural Network (Auto-Net) for Brain Extraction in Magnetic Resonance Imaging
- Iterative Multi-domain Regularized Deep Learning for Anatomical Structure Detection and Segmentation from Ultrasound Images
- Facial Micro-Expression Spotting and Recognition using Time Contrasted Feature with Visual Memory
- Unified Perceptual Parsing for Scene Understanding
- End-to-end Convolutional Network for Saliency Prediction
- PRSeg: A Lightweight Patch Rotate MLP Decoder for Semantic Segmentation
- Object Detection Free Instance Segmentation With Labeling Transformations
- Optical Flow with Semantic Segmentation and Localized Layers
- Neural Contourlet Network for Monocular 360 Depth Estimation
- Enhancing sea ice segmentation in Sentinel-1 images with atrous convolutions
- Cross-domain Human Parsing via Adversarial Feature and Label Adaptation
- Semantic Video Segmentation by Gated Recurrent Flow Propagation
- Interpretable and Globally Optimal Prediction for Textual Grounding using Image Concepts
- Learning Deep Representations for Semantic Image Parsing: a Comprehensive Overview
- Low-latency Perception in Off-Road Dynamical Low Visibility Environments
- Pose2Instance: Harnessing Keypoints for Person Instance Segmentation
- Fast, Exact and Multi-Scale Inference for Semantic Image Segmentation with Deep Gaussian CRFs
- DenseAttentionSeg: Segment Hands from Interacted Objects Using Depth Input
- Learning Sparse High Dimensional Filters: Image Filtering, Dense CRFs and Bilateral Neural Networks
- An End-to-End Approach to Natural Language Object Retrieval via Context-Aware Deep Reinforcement Learning
- Label-Driven Reconstruction for Domain Adaptation in Semantic Segmentation
- PetroSurf3D - A Dataset for high-resolution 3D Surface Segmentation
- Shallow Network Based on Depthwise Over-Parameterized Convolution for Hyperspectral Image Classification
- SFSegNet: Parse Freehand Sketches using Deep Fully Convolutional Networks
- Self-supervised blur detection from synthetically blurred scenes
- Joint Multi-Person Pose Estimation and Semantic Part Segmentation
- Segmenting Transparent Objects in the Wild
- Learning Transferrable Knowledge for Semantic Segmentation with Deep Convolutional Neural Network
- Deep Feature Flow for Video Recognition
- A Comprehensive Study on Colorectal Polyp Segmentation with ResUNet++, Conditional Random Field and Test-Time Augmentation
- MuraNet: Multi-task Floor Plan Recognition with Relation Attention
- Multi-Scale Feature Aggregation by Cross-Scale Pixel-to-Region Relation Operation for Semantic Segmentation
- Feedback Neural Network for Weakly Supervised Geo-Semantic Segmentation
- Material Segmentation of Multi-View Satellite Imagery
- SketchParse : Towards Rich Descriptions for Poorly Drawn Sketches using Multi-Task Hierarchical Deep Networks
- Advances in Medical Image Segmentation: A Comprehensive Survey with a Focus on Lumbar Spine Applications
- End-to-End Training of Hybrid CNN-CRF Models for Stereo
- On the Importance of Visual Context for Data Augmentation in Scene Understanding
- Segmentation method of U-net sheet metal engineering drawing based on CBAM attention mechanism
- Is Faster R-CNN Doing Well for Pedestrian Detection?
- Propagating Asymptotic-Estimated Gradients for Low Bitwidth Quantized Neural Networks
- Dense Recurrent Neural Networks for Scene Labeling
- Deep Structured Scene Parsing by Learning with Image Descriptions
- Surveillance Video Parsing with Single Frame Supervision
- Error Correction for Dense Semantic Image Labeling
- Importance Sampling CAMs for Weakly-Supervised Segmentation
- SegStereo: Exploiting Semantic Information for Disparity Estimation
- Towards High Performance Video Object Detection
- BusyHands: A Hand-Tool Interaction Database for Assembly Tasks Semantic Segmentation
- Multi-scale Attention U-Net (MsAUNet): A Modified U-Net Architecture for Scene Segmentation
- Video Object Segmentation with Joint Re-identification and Attention-Aware Mask Propagation
- SAR-U-Net: squeeze-and-excitation block and atrous spatial pyramid pooling based residual U-Net for automatic liver segmentation in Computed Tomography
- Boundary Corrected Multi-scale Fusion Network for Real-time Semantic Segmentation
- Improving Contrastive Learning by Visualizing Feature Transformation
- Deep Flow-Guided Video Inpainting
- Learning Motion Patterns in Videos
- Exploration of Convolutional Neural Network Architectures for Large Region Map Automation
- Integrated Inference and Learning of Neural Factors in Structural Support Vector Machines
- TriangleNet: Edge Prior Augmented Network for Semantic Segmentation through Cross-Task Consistency
- Fully Convolutional Neural Network for Semantic Segmentation of Anatomical Structure and Pathologies in Colour Fundus Images Associated with Diabetic Retinopathy
- Low-Latency Video Semantic Segmentation
- Bipartite Conditional Random Fields for Panoptic Segmentation
- Multilevel Context Representation for Improving Object Recognition
- Semantic Segmentation with Boundary Neural Fields
- Diverse Sampling for Self-Supervised Learning of Semantic Segmentation
- Open Compound Domain Adaptation
- Saliency guided deep network for weakly-supervised image segmentation
- Learnable Histogram: Statistical Context Features for Deep Neural Networks
- Dynamic Video Segmentation Network
- One-Shot Segmentation in Clutter
- Unsupervised Learning of Important Objects from First-Person Videos
- Nazr-CNN: Fine-Grained Classification of UAV Imagery for Damage Assessment
- Better Image Segmentation by Exploiting Dense Semantic Predictions
- STD2P: RGBD Semantic Segmentation Using Spatio-Temporal Data-Driven Pooling
- Fast Video Object Segmentation With Temporal Aggregation Network and Dynamic Template Matching
- Object Detection via Aspect Ratio and Context Aware Region-based Convolutional Networks
- Defective samples simulation through Neural Style Transfer for automatic surface defect segment
- Feature-Fused Context-Encoding Network for Neuroanatomy Segmentation
- Constrained Convolutional Neural Networks for Weakly Supervised Segmentation
- Scene Parsing via Dense Recurrent Neural Networks with Attentional Selection
- Integrated Deep and Shallow Networks for Salient Object Detection
- Recurrent Multimodal Interaction for Referring Image Segmentation
- Neural Rejuvenation: Improving Deep Network Training by Enhancing Computational Resource Utilization
- Deep Markov Random Field for Image Modeling
- Semantic-Rearrangement-Based Multi-Level Alignment for Domain Generalized Segmentation
- Principled Parallel Mean-Field Inference for Discrete Random Fields
- Framework-agnostic Semantically-aware Global Reasoning for Segmentation
- Feature Selective Networks for Object Detection
- Deep Learning Estimation of Absorbed Dose for Nuclear Medicine Diagnostics
- Adversarial Learning for Image Forensics Deep Matching with Atrous Convolution
- Context Prior for Scene Segmentation
- Unsupervised data augmentation for object detection
- 3D Guided Weakly Supervised Semantic Segmentation
- StuffNet: Using 'Stuff' to Improve Object Detection
- Beyond Planar Symmetry: Modeling human perception of reflection and rotation symmetries in the wild
- What Can Help Pedestrian Detection?
- Learning Dilation Factors for Semantic Segmentation of Street Scenes
- Split-Merge Pooling
- MoE-SPNet: A Mixture-of-Experts Scene Parsing Network
- Using Cross-Model EgoSupervision to Learn Cooperative Basketball Intention
- DV3+HED+: A DCNNs-based Framework to Monitor Temporary Works and ESAs in Railway Construction Project Using VHR Satellite Images
- Context Based Visual Content Verification
- Per-Pixel Feedback for improving Semantic Segmentation
- Compact retail shelf segmentation for mobile deployment
- A De-raining semantic segmentation network for real-time foreground segmentation
- Triplet-based Deep Similarity Learning for Person Re-Identification
- A Multi-Layer Approach to Superpixel-based Higher-order Conditional Random Field for Semantic Image Segmentation