Laplacian Pyramid Reconstruction and Refinement for Semantic Segmentation
arXiv:1605.02264
Abstract
CNN architectures have terrific recognition performance but rely on spatial pooling which makes it difficult to adapt them to tasks that require dense, pixel-accurate labeling. This paper makes two contributions: (1) We demonstrate that while the apparent spatial resolution of convolutional feature maps is low, the high-dimensional feature representation contains significant sub-pixel localization information. (2) We describe a multi-resolution reconstruction architecture based on a Laplacian pyramid that uses skip connections from higher resolution feature maps and multiplicative gating to successively refine segment boundaries reconstructed from lower-resolution maps. This approach yields state-of-the-art semantic segmentation results on the PASCAL VOC and Cityscapes segmentation benchmarks without resorting to more complex random-field inference or instance detection driven architectures.
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Deep Generative Image Models using a Laplacian Pyramid of Adversarial Networks
- The Cityscapes Dataset for Semantic Urban Scene Understanding
- Learning Deconvolution Network for Semantic Segmentation
- BoxSup: Exploiting Bounding Boxes to Supervise Convolutional Networks for Semantic Segmentation
- Pixel-level Encoding and Depth Layering for Instance-level Semantic Labeling
Cited by in corpus (5)
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- ContextNet: Exploring Context and Detail for Semantic Segmentation in Real-time
- A Survey on Deep Learning Methods for Robot Vision
- Deep Learning For Face Recognition: A Critical Analysis
- Deep Inception-Residual Laplacian Pyramid Networks for Accurate Single Image Super-Resolution