Understanding Convolution for Semantic Segmentation
arXiv:1702.08502
Abstract
Recent advances in deep learning, especially deep convolutional neural networks (CNNs), have led to significant improvement over previous semantic segmentation systems. Here we show how to improve pixel-wise semantic segmentation by manipulating convolution-related operations that are of both theoretical and practical value. First, we design dense upsampling convolution (DUC) to generate pixel-level prediction, which is able to capture and decode more detailed information that is generally missing in bilinear upsampling. Second, we propose a hybrid dilated convolution (HDC) framework in the encoding phase. This framework 1) effectively enlarges the receptive fields (RF) of the network to aggregate global information; 2) alleviates what we call the "gridding issue" caused by the standard dilated convolution operation. We evaluate our approaches thoroughly on the Cityscapes dataset, and achieve a state-of-art result of 80.1% mIOU in the test set at the time of submission. We also have achieved state-of-the-art overall on the KITTI road estimation benchmark and the PASCAL VOC2012 segmentation task. Our source code can be found at https://github.com/TuSimple/TuSimple-DUC .
WACV 2018. Updated acknowledgements. Source code: https://github.com/TuSimple/TuSimple-DUC
Cited by in corpus (16)
- Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation
- Dilated Residual Networks
- BiSeNet: Bilateral Segmentation Network for Real-time Semantic Segmentation
- PointSeg: Real-Time Semantic Segmentation Based on 3D LiDAR Point Cloud
- Learning a Discriminative Feature Network for Semantic Segmentation
- Building Footprint Generation Using Improved Generative Adversarial Networks
- ExFuse: Enhancing Feature Fusion for Semantic Segmentation
- MaskLab: Instance Segmentation by Refining Object Detection with Semantic and Direction Features
- Vortex Pooling: Improving Context Representation in Semantic Segmentation
- Scaling Wide Residual Networks for Panoptic Segmentation
- Star Shape Prior in Fully Convolutional Networks for Skin Lesion Segmentation
- Naive-Student: Leveraging Semi-Supervised Learning in Video Sequences for Urban Scene Segmentation
- Benanza: Automatic Benchmark Generation to Compute "Lower-bound" Latency and Inform Optimizations of Deep Learning Models on GPUs
- Scene Understanding Networks for Autonomous Driving based on Around View Monitoring System
- Triply Supervised Decoder Networks for Joint Detection and Segmentation
- ARMA Nets: Expanding Receptive Field for Dense Prediction