Analyzing Modular CNN Architectures for Joint Depth Prediction and Semantic Segmentation
arXiv:1702.08009 · doi:10.1109/ICRA.2017.7989537
Abstract
This paper addresses the task of designing a modular neural network architecture that jointly solves different tasks. As an example we use the tasks of depth estimation and semantic segmentation given a single RGB image. The main focus of this work is to analyze the cross-modality influence between depth and semantic prediction maps on their joint refinement. While most previous works solely focus on measuring improvements in accuracy, we propose a way to quantify the cross-modality influence. We show that there is a relationship between final accuracy and cross-modality influence, although not a simple linear one. Hence a larger cross-modality influence does not necessarily translate into an improved accuracy. We find that a beneficial balance between the cross-modality influences can be achieved by network architecture and conjecture that this relationship can be utilized to understand different network design choices. Towards this end we propose a Convolutional Neural Network (CNN) architecture that fuses the state of the state-of-the-art results for depth estimation and semantic labeling. By balancing the cross-modality influences between depth and semantic prediction, we achieve improved results for both tasks using the NYU-Depth v2 benchmark.
Accepted to ICRA 2017
References in corpus (3)
Cited by in corpus (13)
- PAD-Net: Multi-Tasks Guided Prediction-and-Distillation Network for Simultaneous Depth Estimation and Scene Parsing
- J-MOD: Joint Monocular Obstacle Detection and Depth Estimation
- Wasserstein Distances for Stereo Disparity Estimation
- AuxNet: Auxiliary tasks enhanced Semantic Segmentation for Automated Driving
- Boosting Deep Open World Recognition by Clustering
- RealMonoDepth: Self-Supervised Monocular Depth Estimation for General Scenes
- Towards Real-Time Monocular Depth Estimation for Robotics: A Survey
- A Projected Gradient Descent Method for CRF Inference allowing End-To-End Training of Arbitrary Pairwise Potentials
- Geo-Supervised Visual Depth Prediction
- GAPLE: Generalizable Approaching Policy LEarning for Robotic Object Searching in Indoor Environment
- Learning to Reconstruct and Understand Indoor Scenes from Sparse Views
- Predicting Depth from Semantic Segmentation using Game Engine Dataset
- Split-Merge Pooling