Cascade Residual Learning: A Two-stage Convolutional Neural Network for Stereo Matching
arXiv:1708.09204
Abstract
Leveraging on the recent developments in convolutional neural networks (CNNs), matching dense correspondence from a stereo pair has been cast as a learning problem, with performance exceeding traditional approaches. However, it remains challenging to generate high-quality disparities for the inherently ill-posed regions. To tackle this problem, we propose a novel cascade CNN architecture composing of two stages. The first stage advances the recently proposed DispNet by equipping it with extra up-convolution modules, leading to disparity images with more details. The second stage explicitly rectifies the disparity initialized by the first stage; it couples with the first-stage and generates residual signals across multiple scales. The summation of the outputs from the two stages gives the final disparity. As opposed to directly learning the disparity at the second stage, we show that residual learning provides more effective refinement. Moreover, it also benefits the training of the overall cascade network. Experimentation shows that our cascade residual learning scheme provides state-of-the-art performance for matching stereo correspondence. By the time of the submission of this paper, our method ranks first in the KITTI 2015 stereo benchmark, surpassing the prior works by a noteworthy margin.
Accepted at ICCVW 2017. The first two authors contributed equally to this paper
References in corpus (3)
Cited by in corpus (18)
- Efficient Attention: Attention with Linear Complexities
- Pyramid Stereo Matching Network
- Vehicle Detection of Multi-source Remote Sensing Data Using Active Fine-tuning Network
- Learning for Disparity Estimation through Feature Constancy
- DeepSignals: Predicting Intent of Drivers Through Visual Signals
- Single View Stereo Matching
- Left-Right Comparative Recurrent Model for Stereo Matching
- Learning to Refine Human Pose Estimation
- On the Importance of Stereo for Accurate Depth Estimation: An Efficient Semi-Supervised Deep Neural Network Approach
- Zoom and Learn: Generalizing Deep Stereo Matching to Novel Domains
- NavigationNet: A Large-scale Interactive Indoor Navigation Dataset
- Confidence Inference for Focused Learning in Stereo Matching
- DSR: Direct Self-rectification for Uncalibrated Dual-lens Cameras
- Eliminating the Blind Spot: Adapting 3D Object Detection and Monocular Depth Estimation to 360° Panoramic Imagery
- Multi-scale Cross-form Pyramid Network for Stereo Matching
- Non-destructive three-dimensional measurement of hand vein based on self-supervised network
- Shift Convolution Network for Stereo Matching
- Unsupervised monocular stereo matching