Stereo Matching by Training a Convolutional Neural Network to Compare Image Patches
arXiv:1510.05970
Abstract
We present a method for extracting depth information from a rectified image pair. Our approach focuses on the first stage of many stereo algorithms: the matching cost computation. We approach the problem by learning a similarity measure on small image patches using a convolutional neural network. Training is carried out in a supervised manner by constructing a binary classification data set with examples of similar and dissimilar pairs of patches. We examine two network architectures for this task: one tuned for speed, the other for accuracy. The output of the convolutional neural network is used to initialize the stereo matching cost. A series of post-processing steps follow: cross-based cost aggregation, semiglobal matching, a left-right consistency check, subpixel enhancement, a median filter, and a bilateral filter. We evaluate our method on the KITTI 2012, KITTI 2015, and Middlebury stereo data sets and show that it outperforms other approaches on all three data sets.
References in corpus (2)
Cited by in corpus (42)
- A Survey on Deep Learning Techniques for Stereo-based Depth Estimation
- Continuous 3D Label Stereo Matching using Local Expansion Moves
- Look Wider to Match Image Patches with Convolutional Neural Networks
- PVStereo: Pyramid Voting Module for End-to-End Self-Supervised Stereo Matching
- Improved Multiple-Image-Based Reflection Removal Algorithm Using Deep Neural Networks
- ResDepth: A Deep Residual Prior For 3D Reconstruction From High-resolution Satellite Images
- Neural Markov Random Field for Stereo Matching
- WHU-Stereo: A Challenging Benchmark for Stereo Matching of High-Resolution Satellite Images
- DecomposeMe: Simplifying ConvNets for End-to-End Learning
- Point Set Voting for Partial Point Cloud Analysis
- Simultaneous multi-view instance detection with learned geometric soft-constraints
- Learning Dense Correspondence via 3D-guided Cycle Consistency
- DCVSMNet: Double Cost Volume Stereo Matching Network
- Depth from Monocular Images using a Semi-Parallel Deep Neural Network (SPDNN) Hybrid Architecture
- Adaptive Unimodal Cost Volume Filtering for Deep Stereo Matching
- Artificial Intelligence Enhances the Performance of Chaos-based Wireless Communication
- Robust and accurate depth estimation by fusing LiDAR and Stereo
- Disparity-based Stereo Image Compression with Aligned Cross-View Priors
- Semantic Labeling of Large-Area Geographic Regions Using Multi-View and Multi-Date Satellite Images and Noisy OSM Training Labels
- SOCRATES: A Stereo Camera Trap for Monitoring of Biodiversity
- Visual Depth Mapping from Monocular Images using Recurrent Convolutional Neural Networks
- Multi-tiling Neural Radiance Field (NeRF) -- Geometric Assessment on Large-scale Aerial Datasets
- Dense Semantic Forecasting in Video by Joint Regression of Features and Feature Motion
- YOLOStereo3D: A Step Back to 2D for Efficient Stereo 3D Detection
- SceneEDNet: A Deep Learning Approach for Scene Flow Estimation
- TW-SMNet: Deep Multitask Learning of Tele-Wide Stereo Matching
- A Novel Monocular Disparity Estimation Network with Domain Transformation and Ambiguity Learning
- LeanStereo: A Leaner Backbone based Stereo Network
- An evaluation of Deep Learning based stereo dense matching dataset shift from aerial images and a large scale stereo dataset
- TIDE: Temporally Incremental Disparity Estimation via Pattern Flow in Structured Light System
- A Novel Factor Graph-Based Optimization Technique for Stereo Correspondence Estimation
- DeepSim-Nets: Deep Similarity Networks for Stereo Image Matching
- Distilling Stereo Networks for Performant and Efficient Leaner Networks
- Stereo Matching by Joint Energy Minimization
- Image Patch Matching Using Convolutional Descriptors with Euclidean Distance
- Deep Learning Stereo Vision at the edge
- Detecting Ground Control Points via Convolutional Neural Network for Stereo Matching
- Improving Depth Estimation using Location Information
- Finding Correspondences for Optical Flow and Disparity Estimations using a Sub-pixel Convolution-based Encoder-Decoder Network
- Non-destructive three-dimensional measurement of hand vein based on self-supervised network
- Monocular Depth Estimation with Directional Consistency by Deep Networks
- Shift Convolution Network for Stereo Matching