Predicting Depth, Surface Normals and Semantic Labels with a Common Multi-Scale Convolutional Architecture
arXiv:1411.4734
Abstract
In this paper we address three different computer vision tasks using a single basic architecture: depth prediction, surface normal estimation, and semantic labeling. We use a multiscale convolutional network that is able to adapt easily to each task using only small modifications, regressing from the input image to the output map directly. Our method progressively refines predictions using a sequence of scales, and captures many image details without any superpixels or low-level segmentation. We achieve state-of-the-art performance on benchmarks for all three tasks.
References in corpus (12)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- Going Deeper with Convolutions
- OverFeat: Integrated Recognition, Localization and Detection using Convolutional Networks
- Computing the Stereo Matching Cost with a Convolutional Neural Network
- Fully Convolutional Networks for Semantic Segmentation
- Indoor Semantic Segmentation using depth information
- Designing Deep Networks for Surface Normal Estimation
- Deep and Wide Multiscale Recursive Networks for Robust Image Labeling
- Deep Convolutional Neural Fields for Depth Estimation from a Single Image
- Coupled Depth Learning
Cited by in corpus (76)
- Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs
- What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?
- Conditional Random Fields as Recurrent Neural Networks
- Learning from Synthetic Humans
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- Weakly- and Semi-Supervised Learning of a DCNN for Semantic Image Segmentation
- A Survey on Deep Learning-based Architectures for Semantic Segmentation on 2D images
- ImageNet pre-trained models with batch normalization
- A simple yet effective baseline for 3d human pose estimation
- PixelNet: Representation of the pixels, by the pixels, and for the pixels
- Do We Really Need to Collect Millions of Faces for Effective Face Recognition?
- CNN-SLAM: Real-time dense monocular SLAM with learned depth prediction
- PAD-Net: Multi-Tasks Guided Prediction-and-Distillation Network for Simultaneous Depth Estimation and Scene Parsing
- Improved Adversarial Systems for 3D Object Generation and Reconstruction
- Recurrent Slice Networks for 3D Segmentation of Point Clouds
- Structured Attention Guided Convolutional Neural Fields for Monocular Depth Estimation
- 3D Shape Reconstruction from Sketches via Multi-view Convolutional Networks
- ZM-Net: Real-time Zero-shot Image Manipulation Network
- Joint Object and Part Segmentation using Deep Learned Potentials
- Joint Sequence Learning and Cross-Modality Convolution for 3D Biomedical Segmentation
- Attention-based Context Aggregation Network for Monocular Depth Estimation
- Deep Depth Completion of a Single RGB-D Image
- LEGO: Learning Edge with Geometry all at Once by Watching Videos
- 3D Photography using Context-aware Layered Depth Inpainting
- Unsupervised Learning of Monocular Depth Estimation with Bundle Adjustment, Super-Resolution and Clip Loss
- 3D Shape Induction from 2D Views of Multiple Objects
- Scene Memory Transformer for Embodied Agents in Long-Horizon Tasks
- Borrowing Treasures from the Wealthy: Deep Transfer Learning through Selective Joint Fine-tuning
- PlaneNet: Piece-wise Planar Reconstruction from a Single RGB Image
- Learning to Synthesize a 4D RGBD Light Field from a Single Image
- Single View Stereo Matching
- Geometry-aware Deep Network for Single-Image Novel View Synthesis
- Joint Prediction of Depths, Normals and Surface Curvature from RGB Images using CNNs
- Deep Outdoor Illumination Estimation
- DSAC - Differentiable RANSAC for Camera Localization
- Estimated Depth Map Helps Image Classification
- Im2Struct: Recovering 3D Shape Structure from a Single RGB Image
- Generalizing multistain immunohistochemistry tissue segmentation using one-shot color deconvolution deep neural networks
- Every Pixel Counts: Unsupervised Geometry Learning with Holistic 3D Motion Understanding
- Depth Information Guided Crowd Counting for Complex Crowd Scenes
- Weakly Supervised Object Localization Using Things and Stuff Transfer
- Salient Region Segmentation
- Adaptive Weighting Multi-Field-of-View CNN for Semantic Segmentation in Pathology
- FusionLane: Multi-Sensor Fusion for Lane Marking Semantic Segmentation Using Deep Neural Networks
- Semantic Object Parsing with Graph LSTM
- CAM-Convs: Camera-Aware Multi-Scale Convolutions for Single-View Depth
- DeLS-3D: Deep Localization and Segmentation with a 3D Semantic Map
- Error Correction for Dense Semantic Image Labeling
- PixelNN: Example-based Image Synthesis
- Traffic Density Estimation using a Convolutional Neural Network
- Pano2CAD: Room Layout From A Single Panorama Image
- Value of Temporal Dynamics Information in Driving Scene Segmentation
- Radar Emitter Classification with Attribute-specific Recurrent Neural Networks
- Layer-structured 3D Scene Inference via View Synthesis
- DF-Net: Unsupervised Joint Learning of Depth and Flow using Cross-Task Consistency
- Physically-Based Rendering for Indoor Scene Understanding Using Convolutional Neural Networks
- Cross-Domain Self-supervised Multi-task Feature Learning using Synthetic Imagery
- SeGAN: Segmenting and Generating the Invisible
- FrameNet: Learning Local Canonical Frames of 3D Surfaces from a Single RGB Image
- Binge Watching: Scaling Affordance Learning from Sitcoms
- Normal Assisted Stereo Depth Estimation
- Detect, Replace, Refine: Deep Structured Prediction For Pixel Wise Labeling
- Instance-Level Salient Object Segmentation
- 3D Neighborhood Convolution: Learning Depth-Aware Features for RGB-D and RGB Semantic Segmentation
- Single Image Depth Estimation Trained via Depth from Defocus Cues
- Deep Reflectance Maps
- Shape from Shading through Shape Evolution
- Critical Contours: An Invariant Linking Image Flow with Salient Surface Organization
- VLASE: Vehicle Localization by Aggregating Semantic Edges
- Improving task-specific representation via 1M unlabelled images without any extra knowledge
- Depth Assisted Full Resolution Network for Single Image-based View Synthesis
- Split-Merge Pooling
- DeepErase: Weakly Supervised Ink Artifact Removal in Document Text Images
- Authoring image decompositions with generative models
- Towards Comprehensive Monocular Depth Estimation: Multiple Heads Are Better Than One
- Monocular Depth Estimation with Directional Consistency by Deep Networks