Tree-structured Kronecker Convolutional Network for Semantic Segmentation
arXiv:1812.04945
Abstract
Most existing semantic segmentation methods employ atrous convolution to enlarge the receptive field of filters, but neglect partial information. To tackle this issue, we firstly propose a novel Kronecker convolution which adopts Kronecker product to expand the standard convolutional kernel for taking into account the partial feature neglected by atrous convolutions. Therefore, it can capture partial information and enlarge the receptive field of filters simultaneously without introducing extra parameters. Secondly, we propose Tree-structured Feature Aggregation (TFA) module which follows a recursive rule to expand and forms a hierarchical structure. Thus, it can naturally learn representations of multi-scale objects and encode hierarchical contextual information in complex scenes. Finally, we design Tree-structured Kronecker Convolutional Networks (TKCN) which employs Kronecker convolution and TFA module. Extensive experiments on three datasets, PASCAL VOC 2012, PASCAL-Context and Cityscapes, verify the effectiveness of our proposed approach. We make the code and the trained model publicly available at https://github.com/wutianyiRosun/TKCN.
Code: https://github.com/wutianyiRosun/TKCN
References in corpus (14)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- The Cityscapes Dataset for Semantic Urban Scene Understanding
- Learning Deconvolution Network for Semantic Segmentation
- Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- Context Encoding for Semantic Segmentation
- Learning a Discriminative Feature Network for Semantic Segmentation
- CGNet: A Light-weight Context Guided Network for Semantic Segmentation
- COCO-Stuff: Thing and Stuff Classes in Context
- Not All Pixels Are Equal: Difficulty-aware Semantic Segmentation via Deep Layer Cascade
- Exploring Context with Deep Structured models for Semantic Segmentation
- FoveaNet: Perspective-aware Urban Scene Parsing