Combining the Best of Convolutional Layers and Recurrent Layers: A Hybrid Network for Semantic Segmentation
arXiv:1603.04871
Abstract
State-of-the-art results of semantic segmentation are established by Fully Convolutional neural Networks (FCNs). FCNs rely on cascaded convolutional and pooling layers to gradually enlarge the receptive fields of neurons, resulting in an indirect way of modeling the distant contextual dependence. In this work, we advocate the use of spatially recurrent layers (i.e. ReNet layers) which directly capture global contexts and lead to improved feature representations. We demonstrate the effectiveness of ReNet layers by building a Naive deep ReNet (N-ReNet), which achieves competitive performance on Stanford Background dataset. Furthermore, we integrate ReNet layers with FCNs, and develop a novel Hybrid deep ReNet (H-ReNet). It enjoys a few remarkable properties, including full-image receptive fields, end-to-end training, and efficient network execution. On the PASCAL VOC 2012 benchmark, the H-ReNet improves the results of state-of-the-art approaches Piecewise, CRFasRNN and DeepParsing by 3.6%, 2.3% and 0.2%, respectively, and achieves the highest IoUs for 13 out of the 20 object classes.
14 pages
References in corpus (8)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Recurrent Neural Network Regularization
- Going Deeper with Convolutions
- Learning Deconvolution Network for Semantic Segmentation
- Fully Connected Deep Structured Networks
- Semantic Image Segmentation via Deep Parsing Network
- BoxSup: Exploiting Bounding Boxes to Supervise Convolutional Networks for Semantic Segmentation
- Inside-Outside Net: Detecting Objects in Context with Skip Pooling and Recurrent Neural Networks
Cited by in corpus (11)
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- A Deep Spatial Contextual Long-term Recurrent Convolutional Network for Saliency Detection
- Semantic Correlation Promoted Shape-Variant Context for Segmentation
- Contrastive Proposal Extension with LSTM Network for Weakly Supervised Object Detection
- Partial Labeled Gastric Tumor Segmentation via patch-based Reiterative Learning
- Beyond Forward Shortcuts: Fully Convolutional Master-Slave Networks (MSNets) with Backward Skip Connections for Semantic Segmentation
- Scene Labeling using Gated Recurrent Units with Explicit Long Range Conditioning
- Semantic Segmentation of Earth Observation Data Using Multimodal and Multi-scale Deep Networks
- Cavs: A Vertex-centric Programming Interface for Dynamic Neural Networks
- Self-Attentive Multi-Layer Aggregation with Feature Recalibration and Normalization for End-to-End Speaker Verification System
- RethNet: Object-by-Object Learning for Detecting Facial Skin Problems