Exploring Context with Deep Structured models for Semantic Segmentation
arXiv:1603.03183
Abstract
State-of-the-art semantic image segmentation methods are mostly based on training deep convolutional neural networks (CNNs). In this work, we proffer to improve semantic segmentation with the use of contextual information. In particular, we explore `patch-patch' context and `patch-background' context in deep CNNs. We formulate deep structured models by combining CNNs and Conditional Random Fields (CRFs) for learning the patch-patch context between image regions. Specifically, we formulate CNN-based pairwise potential functions to capture semantic correlations between neighboring patches. Efficient piecewise training of the proposed deep structured model is then applied in order to avoid repeated expensive CRF inference during the course of back propagation. For capturing the patch-background context, we show that a network design with traditional multi-scale image inputs and sliding pyramid pooling is very effective for improving performance. We perform comprehensive evaluation of the proposed method. We achieve new state-of-the-art performance on a number of challenging semantic segmentation datasets including , -, , -, -, -, and datasets. Particularly, we report an intersection-over-union score of on the - dataset.
16 pages. Accepted to IEEE T. Pattern Analysis & Machine Intelligence, 2017. Extended version of arXiv:1504.01013
References in corpus (10)
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- The Cityscapes Dataset for Semantic Urban Scene Understanding
- Learning Deconvolution Network for Semantic Segmentation
- Practical recommendations for gradient-based training of deep architectures
- Fully Connected Deep Structured Networks
- Simultaneous Detection and Segmentation
- BoxSup: Exploiting Bounding Boxes to Supervise Convolutional Networks for Semantic Segmentation
- Piecewise Training for Undirected Models
- Closed-Form Training of Conditional Random Fields for Large Scale Image Segmentation
Cited by in corpus (9)
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- RedNet: Residual Encoder-Decoder Network for indoor RGB-D Semantic Segmentation
- Simultaneous Traffic Sign Detection and Boundary Estimation using Convolutional Neural Network
- Bridging Category-level and Instance-level Semantic Image Segmentation
- Multi-View Deep Learning for Consistent Semantic Mapping with RGB-D Cameras
- FoveaNet: Perspective-aware Urban Scene Parsing
- Tree-structured Kronecker Convolutional Network for Semantic Segmentation
- Visual-Inertial-Semantic Scene Representation for 3-D Object Detection
- Learning Multi-level Region Consistency with Dense Multi-label Networks for Semantic Segmentation