Efficient piecewise training of deep structured models for semantic segmentation
arXiv:1504.01013
Abstract
Recent advances in semantic image segmentation have mostly been achieved by training deep convolutional neural networks (CNNs). We show how to improve semantic segmentation through the use of contextual information; specifically, we explore `patch-patch' context between image regions, and `patch-background' context. For learning from the patch-patch context, we formulate Conditional Random Fields (CRFs) with CNN-based pairwise potential functions to capture semantic correlations between neighboring patches. Efficient piecewise training of the proposed deep structured model is then applied to avoid repeated expensive CRF inference for back propagation. For capturing the patch-background context, we show that a network design with traditional multi-scale image input and sliding pyramid pooling is effective for improving performance. Our experimental results set new state-of-the-art performance on a number of popular semantic segmentation datasets, including NYUDv2, PASCAL VOC 2012, PASCAL-Context, and SIFT-flow. In particular, we achieve an intersection-over-union score of 78.0 on the challenging PASCAL VOC 2012 dataset.
Appearing in IEEE Conf. Computer Vision and Pattern Recognition (CVPR) 2016
References in corpus (11)
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- MatConvNet - Convolutional Neural Networks for MATLAB
- Learning Deconvolution Network for Semantic Segmentation
- Convolutional Feature Masking for Joint Object and Stuff Segmentation
- Fully Connected Deep Structured Networks
- Simultaneous Detection and Segmentation
- Semantic Image Segmentation via Deep Parsing Network
- BoxSup: Exploiting Bounding Boxes to Supervise Convolutional Networks for Semantic Segmentation
- Piecewise Training for Undirected Models
- Closed-Form Training of Conditional Random Fields for Large Scale Image Segmentation
Cited by in corpus (15)
- Conditional Random Fields as Recurrent Neural Networks
- Semantic Image Segmentation via Deep Parsing Network
- Deep Learning Convolutional Networks for Multiphoton Microscopy Vasculature Segmentation
- The Stixel world: A medium-level representation of traffic scenes
- Semantic Image Segmentation with Task-Specific Edge Detection Using CNNs and a Discriminatively Trained Domain Transform
- Deeply Learning the Messages in Message Passing Inference
- Joint Object and Part Segmentation using Deep Learned Potentials
- Multi-modal land cover mapping of remote sensing images using pyramid attention and gated fusion networks
- Discriminative Training of Deep Fully-connected Continuous CRF with Task-specific Loss
- Attention-based Context Aggregation Network for Monocular Depth Estimation
- ORDNet: Capturing Omni-Range Dependencies for Scene Parsing
- Deep Learning Markov Random Field for Semantic Segmentation
- FusionLane: Multi-Sensor Fusion for Lane Marking Semantic Segmentation Using Deep Neural Networks
- From Selective Deep Convolutional Features to Compact Binary Representations for Image Retrieval
- Pushing the Limits of Deep CNNs for Pedestrian Detection