Semantic Image Segmentation via Deep Parsing Network
arXiv:1509.02634
Abstract
This paper addresses semantic image segmentation by incorporating rich information into Markov Random Field (MRF), including high-order relations and mixture of label contexts. Unlike previous works that optimized MRFs using iterative algorithm, we solve MRF by proposing a Convolutional Neural Network (CNN), namely Deep Parsing Network (DPN), which enables deterministic end-to-end computation in a single forward pass. Specifically, DPN extends a contemporary CNN architecture to model unary terms and additional layers are carefully devised to approximate the mean field algorithm (MF) for pairwise terms. It has several appealing properties. First, different from the recent works that combined CNN and MRF, where many iterations of MF were required for each training image during back-propagation, DPN is able to achieve high performance by approximating one iteration of MF. Second, DPN represents various types of pairwise terms, making many existing works as its special cases. Third, DPN makes MF easier to be parallelized and speeded up in Graphical Processing Unit (GPU). DPN is thoroughly evaluated on the PASCAL VOC 2012 dataset, where a single DPN model yields a new state-of-the-art segmentation accuracy.
To appear in International Conference on Computer Vision (ICCV) 2015
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Distilling the Knowledge in a Neural Network
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Fully Convolutional Networks for Semantic Segmentation
- Deep Learning Face Representation by Joint Identification-Verification
- cuDNN: Efficient Primitives for Deep Learning
- Fully Connected Deep Structured Networks
Cited by in corpus (83)
- Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- FCNs in the Wild: Pixel-level Adversarial and Constraint-based Adaptation
- Pyramid Scene Parsing Network
- Pyramid Attention Network for Semantic Segmentation
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- Dual Attention Network for Scene Segmentation
- Training convolutional neural networks to estimate turbulent sub-grid scale reaction rates
- Context Encoding for Semantic Segmentation
- ICNet for Real-Time Semantic Segmentation on High-Resolution Images
- Learning a Discriminative Feature Network for Semantic Segmentation
- Stacked Deconvolutional Network for Semantic Segmentation
- Semantic Facial Expression Editing using Autoencoded Flow
- Video Object Segmentation with Re-identification
- RefineNet: Multi-Path Refinement Networks for High-Resolution Semantic Segmentation
- PAD-Net: Multi-Tasks Guided Prediction-and-Distillation Network for Simultaneous Depth Estimation and Scene Parsing
- Pushing the Boundaries of Boundary Detection using Deep Learning
- The Stixel world: A medium-level representation of traffic scenes
- Combining the Best of Convolutional Layers and Recurrent Layers: A Hybrid Network for Semantic Segmentation
- Attention guided global enhancement and local refinement network for semantic segmentation
- Deep Interactive Object Selection
- MaskLab: Instance Segmentation by Refining Object Detection with Semantic and Direction Features
- Semantic Image Segmentation with Task-Specific Edge Detection Using CNNs and a Discriminatively Trained Domain Transform
- Pixel-level Encoding and Depth Layering for Instance-level Semantic Labeling
- Laplacian Pyramid Reconstruction and Refinement for Semantic Segmentation
- Semantic Object Parsing with Local-Global Long Short-Term Memory
- Weakly Supervised Semantic Segmentation using Web-Crawled Videos
- InstanceCut: from Edges to Instances with MultiCut
- Improving Fully Convolution Network for Semantic Segmentation
- DCNAS: Densely Connected Neural Architecture Search for Semantic Image Segmentation
- Knowledge Adaptation for Efficient Semantic Segmentation
- Evolutionary Synthesis of Deep Neural Networks via Synaptic Cluster-driven Genetic Encoding
- Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes
- Scene Parsing with Global Context Embedding
- Adaptive Affinity Fields for Semantic Segmentation
- Semantic Video Segmentation by Gated Recurrent Flow Propagation
- Unsupervised Total Variation Loss for Semi-supervised Deep Learning of Semantic Segmentation
- Exploring Context with Deep Structured models for Semantic Segmentation
- Zoom Better to See Clearer: Human and Object Parsing with Hierarchical Auto-Zoom Net
- Learning Deep Representations for Semantic Image Parsing: a Comprehensive Overview
- Fast, Exact and Multi-Scale Inference for Semantic Image Segmentation with Deep Gaussian CRFs
- Label-Driven Reconstruction for Domain Adaptation in Semantic Segmentation
- Tensor Low-Rank Reconstruction for Semantic Segmentation
- Recent Advances in the Applications of Convolutional Neural Networks to Medical Image Contour Detection
- Reversible Recursive Instance-level Object Segmentation
- Semi-supervised Domain Adaptation based on Dual-level Domain Mixing for Semantic Segmentation
- CARAFE++: Unified Content-Aware ReAssembly of FEatures
- Fast Semantic Image Segmentation with High Order Context and Guided Filtering
- End-to-End Training of Hybrid CNN-CRF Models for Stereo
- Surveillance Video Parsing with Single Frame Supervision
- Learning to Predict Context-adaptive Convolution for Semantic Segmentation
- Dense Recurrent Neural Networks for Scene Labeling
- Semantic Object Parsing with Graph LSTM
- Deep Structured Scene Parsing by Learning with Image Descriptions
- Deep Convolutional Neural Networks with Spatial Regularization, Volume and Star-shape Priori for Image Segmentation
- Mix-and-Match Tuning for Self-Supervised Semantic Segmentation
- Pseudo Mask Augmented Object Detection
- Multi-scale Attention U-Net (MsAUNet): A Modified U-Net Architecture for Scene Segmentation
- A Projected Gradient Descent Method for CRF Inference allowing End-To-End Training of Arbitrary Pairwise Potentials
- Superpixel Convolutional Networks using Bilateral Inceptions
- Triply Supervised Decoder Networks for Joint Detection and Segmentation
- ScribbleSup: Scribble-Supervised Convolutional Networks for Semantic Segmentation
- SeGAN: Segmenting and Generating the Invisible
- Scene Labeling using Gated Recurrent Units with Explicit Long Range Conditioning
- Hard Pixel Mining for Depth Privileged Semantic Segmentation
- Low-Latency Video Semantic Segmentation
- Bipartite Conditional Random Fields for Panoptic Segmentation
- High-Quality Correspondence and Segmentation Estimation for Dual-Lens Smart-Phone Portraits
- Learning Deep Representations for Scene Labeling with Semantic Context Guided Supervision
- STD2P: RGBD Semantic Segmentation Using Spatio-Temporal Data-Driven Pooling
- Improved Hard Example Mining by Discovering Attribute-based Hard Person Identity
- Dense CNN Learning with Equivalent Mappings
- Learnable Histogram: Statistical Context Features for Deep Neural Networks
- Multi-Level Contextual Network for Biomedical Image Segmentation
- Phase Contrast Microscopy Cell PopulationSegmentation: A Survey
- Deep, Dense, and Low-Rank Gaussian Conditional Random Fields
- Progressively Diffused Networks for Semantic Image Segmentation
- Robust and Efficient Graph Correspondence Transfer for Person Re-identification
- Scene Parsing via Dense Recurrent Neural Networks with Attentional Selection
- RethNet: Object-by-Object Learning for Detecting Facial Skin Problems
- Affinity Derivation and Graph Merge for Instance Segmentation
- A Spatially Constrained Deep Convolutional Neural Network for Nerve Fiber Segmentation in Corneal Confocal Microscopic Images using Inaccurate Annotations
- LSTM-CF: Unifying Context Modeling and Fusion with LSTMs for RGB-D Scene Labeling