Scene Parsing via Dense Recurrent Neural Networks with Attentional Selection
arXiv:1811.04778
Abstract
Recurrent neural networks (RNNs) have shown the ability to improve scene parsing through capturing long-range dependencies among image units. In this paper, we propose dense RNNs for scene labeling by exploring various long-range semantic dependencies among image units. Different from existing RNN based approaches, our dense RNNs are able to capture richer contextual dependencies for each image unit by enabling immediate connections between each pair of image units, which significantly enhances their discriminative power. Besides, to select relevant dependencies and meanwhile to restrain irrelevant ones for each unit from dense connections, we introduce an attention model into dense RNNs. The attention model allows automatically assigning more importance to helpful dependencies while less weight to unconcerned dependencies. Integrating with convolutional neural networks (CNNs), we develop an end-to-end scene labeling system. Extensive experiments on three large-scale benchmarks demonstrate that the proposed approach can improve the baselines by large margins and outperform other state-of-the-art algorithms.
10 pages. arXiv admin note: substantial text overlap with arXiv:1801.06831
References in corpus (14)
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Hierarchical Question-Image Co-Attention for Visual Question Answering
- The Cityscapes Dataset for Semantic Urban Scene Understanding
- Fully Convolutional Networks for Semantic Segmentation
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- Learning Deconvolution Network for Semantic Segmentation
- Pyramid Scene Parsing Network
- BoxSup: Exploiting Bounding Boxes to Supervise Convolutional Networks for Semantic Segmentation
- Context Encoding for Semantic Segmentation
- Multi-Context Attention for Human Pose Estimation
- PixelNet: Representation of the pixels, by the pixels, and for the pixels
- Scene Parsing with Global Context Embedding
- SANet: Structure-Aware Network for Visual Tracking
- Video Scene Parsing with Predictive Feature Learning