RedNet: Residual Encoder-Decoder Network for indoor RGB-D Semantic Segmentation
arXiv:1806.01054
Abstract
Indoor semantic segmentation has always been a difficult task in computer vision. In this paper, we propose an RGB-D residual encoder-decoder architecture, named RedNet, for indoor RGB-D semantic segmentation. In RedNet, the residual module is applied to both the encoder and decoder as the basic building block, and the skip-connection is used to bypass the spatial feature between the encoder and decoder. In order to incorporate the depth information of the scene, a fusion structure is constructed, which makes inference on RGB image and depth image separately, and fuses their features over several layers. In order to efficiently optimize the network's parameters, we propose a `pyramid supervision' training scheme, which applies supervised learning over different layers in the decoder, to cope with the problem of gradients vanishing. Experiment results show that the proposed RedNet(ResNet-50) achieves a state-of-the-art mIoU accuracy of 47.8% on the SUN RGB-D benchmark dataset.
References in corpus (11)
- Wide Residual Networks
- The Cityscapes Dataset for Semantic Urban Scene Understanding
- Learning Deconvolution Network for Semantic Segmentation
- Residual Networks Behave Like Ensembles of Relatively Shallow Networks
- Indoor Semantic Segmentation using depth information
- FusionNet: A deep fully residual convolutional neural network for image segmentation in connectomics
- Dilated Residual Networks
- The Importance of Skip Connections in Biomedical Image Segmentation
- Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes
- Exploring Context with Deep Structured models for Semantic Segmentation
- Incorporating Depth into both CNN and CRF for Indoor Semantic Segmentation
Cited by in corpus (22)
- L3MVN: Leveraging Large Language Models for Visual Target Navigation
- RGB-D And Thermal Sensor Fusion: A Systematic Literature Review
- WayFAST: Navigation with Predictive Traversability in the Field
- RFBNet: Deep Multimodal Networks with Residual Fusion Blocks for RGB-D Semantic Segmentation
- What Makes Multi-modal Learning Better than Single (Provably)
- Frontier Semantic Exploration for Visual Target Navigation
- ACNet: Attention Based Network to Exploit Complementary Features for RGBD Semantic Segmentation
- Improving Multi-Modal Learning with Uni-Modal Teachers
- Revisiting Single Image Depth Estimation: Toward Higher Resolution Maps with Accurate Object Boundaries
- Efficient Multi-Task Scene Analysis with RGB-D Transformers
- Auxiliary Tasks and Exploration Enable ObjectNav
- Global-Local Propagation Network for RGB-D Semantic Segmentation
- Implicit Obstacle Map-driven Indoor Navigation Model for Robust Obstacle Avoidance
- Real-time Fusion Network for RGB-D Semantic Segmentation Incorporating Unexpected Obstacle Detection for Road-driving Images
- Leveraging Large Language Model-based Room-Object Relationships Knowledge for Enhancing Multimodal-Input Object Goal Navigation
- Small Obstacle Avoidance Based on RGB-D Semantic Segmentation
- Residual Encoder-Decoder Network for Deep Subspace Clustering
- Image fusion using symmetric skip autoencodervia an Adversarial Regulariser
- Variation Network: Learning High-level Attributes for Controlled Input Manipulation
- Deep feature selection-and-fusion for RGB-D semantic segmentation
- How semantic and geometric information mutually reinforce each other in ToF object localization
- Deep Iteration Assisted by Multi-level Obey-pixel Network Discriminator (DIAMOND) for Medical Image Recovery