LEGO: Learning Edge with Geometry all at Once by Watching Videos
arXiv:1803.05648
Abstract
Learning to estimate 3D geometry in a single image by watching unlabeled videos via deep convolutional network is attracting significant attention. In this paper, we introduce a "3D as-smooth-as-possible (3D-ASAP)" prior inside the pipeline, which enables joint estimation of edges and 3D scene, yielding results with significant improvement in accuracy for fine detailed structures. Specifically, we define the 3D-ASAP prior by requiring that any two points recovered in 3D from an image should lie on an existing planar surface if no other cues provided. We design an unsupervised framework that Learns Edges and Geometry (depth, normal) all at Once (LEGO). The predicted edges are embedded into depth and surface normal smoothness terms, where pixels without edges in-between are constrained to satisfy the prior. In our framework, the predicted depths, normals and edges are forced to be consistent all the time. We conduct experiments on KITTI to evaluate our estimated geometry and CityScapes to perform edge evaluation. We show that in all of the tasks, i.e.depth, normal and edge, our algorithm vastly outperforms other state-of-the-art (SOTA) algorithms, demonstrating the benefits of our approach.
Accepted to CVPR 2018 as spotlight; Camera ready plus supplementary material. Code will come
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Rethinking Atrous Convolution for Semantic Image Segmentation
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Depth Map Prediction from a Single Image using a Multi-Scale Deep Network
- SfM-Net: Learning of Structure and Motion from Video
- Unsupervised Learning of Depth and Ego-Motion from Video
- Unsupervised Learning of Geometry with Edge-aware Depth-Normal Consistency
- Semi-Supervised Deep Learning for Monocular Depth Map Prediction
- Designing Deep Networks for Surface Normal Estimation
Cited by in corpus (5)
- Autonomous Driving with Deep Learning: A Survey of State-of-Art Technologies
- Monocular Depth Estimation with Self-supervised Instance Adaptation
- Unsupervised Learning of Monocular Depth Estimation with Bundle Adjustment, Super-Resolution and Clip Loss
- Advancing Self-supervised Monocular Depth Learning with Sparse LiDAR
- Unsupervised High-Resolution Depth Learning From Videos With Dual Networks