Computing the Stereo Matching Cost with a Convolutional Neural Network
arXiv:1409.4326 · doi:10.1109/CVPR.2015.7298767
Abstract
We present a method for extracting depth information from a rectified image pair. We train a convolutional neural network to predict how well two image patches match and use it to compute the stereo matching cost. The cost is refined by cross-based cost aggregation and semiglobal matching, followed by a left-right consistency check to eliminate errors in the occluded regions. Our stereo method achieves an error rate of 2.61 % on the KITTI stereo dataset and is currently (August 2014) the top performing method on this dataset.
Conference on Computer Vision and Pattern Recognition (CVPR), June 2015
Cited by in corpus (113)
- Deep learning in remote sensing: a review
- Stereo Matching by Training a Convolutional Neural Network to Compare Image Patches
- DeMoN: Depth and Motion Network for Learning Monocular Stereo
- FlowNet: Learning Optical Flow with Convolutional Networks
- SurfaceNet: An End-to-end 3D Neural Network for Multiview Stereopsis
- A Survey on Deep Learning Techniques for Stereo-based Depth Estimation
- Unsupervised CNN for Single View Depth Estimation: Geometry to the Rescue
- Systematic evaluation of CNN advances on the ImageNet
- End-to-End Learning of Geometry and Context for Deep Stereo Regression
- Learning to Compare Image Patches via Convolutional Neural Networks
- Origami: A 803 GOp/s/W Convolutional Network Accelerator
- Embedded real-time stereo estimation via Semi-Global Matching on the GPU
- Siamese Instance Search for Tracking
- End-to-end Learning of Multi-sensor 3D Tracking by Detection
- DeepToF: Off-the-Shelf Real-Time Correction of Multipath Interference in Time-of-Flight Imaging
- Predicting Depth, Surface Normals and Semantic Labels with a Common Multi-Scale Convolutional Architecture
- LIFT: Learned Invariant Feature Transform
- Generic 3D Representation via Pose Estimation and Matching
- Cascade Cost Volume for High-Resolution Multi-View Stereo and Stereo Matching
- Learning Deep Structure-Preserving Image-Text Embeddings
- Cross-Domain Image Matching with Deep Feature Maps
- GA-Net: Guided Aggregation Net for End-to-end Stereo Matching
- Occlusion-Aware Depth Estimation with Adaptive Normal Constraints
- A Survey on RGB-D Datasets
- Force Estimation from OCT Volumes using 3D CNNs
- Learning Analysis-by-Synthesis for 6D Pose Estimation in RGB-D Images
- AANet: Adaptive Aggregation Network for Efficient Stereo Matching
- Computer Stereo Vision for Autonomous Driving
- Sparse-to-Continuous: Enhancing Monocular Depth Estimation using Occupancy Maps
- A Survey on Deep Learning Architectures for Image-based Depth Reconstruction
- A Large RGB-D Dataset for Semi-supervised Monocular Depth Estimation
- Siamese Network for RGB-D Salient Object Detection and Beyond
- DCVSMNet: Double Cost Volume Stereo Matching Network
- Change Detection between Multimodal Remote Sensing Data Using Siamese CNN
- Scene Flow Estimation: A Survey
- Learning Two-Branch Neural Networks for Image-Text Matching Tasks
- Deep Stereo Matching with Explicit Cost Aggregation Sub-Architecture
- Deep Nets: What have they ever done for Vision?
- Learning Covariant Feature Detectors
- DIML/CVL RGB-D Dataset: 2M RGB-D Images of Natural Indoor and Outdoor Scenes
- RAFT-Stereo: Multilevel Recurrent Field Transforms for Stereo Matching
- Learning the Matching Function
- Learning Dense Convolutional Embeddings for Semantic Segmentation
- Deep Multi-View Stereo gone wild
- FashionNet: Personalized Outfit Recommendation with Deep Neural Network
- Geometry-aware Deep Network for Single-Image Novel View Synthesis
- SimNet: Enabling Robust Unknown Object Manipulation from Pure Synthetic Data via Stereo
- AdaStereo: A Simple and Efficient Approach for Adaptive Stereo Matching
- StereoNet: Guided Hierarchical Refinement for Real-Time Edge-Aware Depth Prediction
- EdgeStereo: An Effective Multi-Task Learning Network for Stereo Matching and Edge Detection
- An In-Depth Analysis of Visual Tracking with Siamese Neural Networks
- Left-Right Comparative Recurrent Model for Stereo Matching
- A Learning-based Framework for Hybrid Depth-from-Defocus and Stereo Matching
- Improved Stereo Matching with Constant Highway Networks and Reflective Confidence Learning
- Recurrent Filter Learning for Visual Tracking
- Learning Video Representations from Correspondence Proposals
- ES-Net: An Efficient Stereo Matching Network
- End-to-end depth from motion with stabilized monocular videos
- UnrealStereo: Controlling Hazardous Factors to Analyze Stereo Vision
- ActiveStereoNet: End-to-End Self-Supervised Learning for Active Stereo Systems
- Video2Shop: Exact Matching Clothes in Videos to Online Shopping Images
- Group-wise Correlation Stereo Network
- Deep Rigid Instance Scene Flow
- Parallax Attention for Unsupervised Stereo Correspondence Learning
- Learning Inverse Depth Regression for Multi-View Stereo with Correlation Cost Volume
- DAVANet: Stereo Deblurring with View Aggregation
- MonSter: Awakening the Mono in Stereo
- Domain-invariant Stereo Matching Networks
- Real-Time Dense Stereo Embedded in A UAV for Road Inspection
- Confidence driven TGV fusion
- Large-Scale 3D Scene Classification With Multi-View Volumetric CNN
- Quick and energy-efficient Bayesian computing of binocular disparity using stochastic digital signals
- Stereo Vision Based Single-Shot 6D Object Pose Estimation for Bin-Picking by a Robot Manipulator
- MVS2D: Efficient Multi-view Stereo via Attention-Driven 2D Convolutions
- A Comparison of Stereo-Matching Cost between Convolutional Neural Network and Census for Satellite Images
- Deep Stereo Matching with Dense CRF Priors
- Bilateral Grid Learning for Stereo Matching Networks
- Computational Eco-Systems for Handwritten Digits Recognition
- Attention Aware Cost Volume Pyramid Based Multi-view Stereo Network for 3D Reconstruction
- AutoScaler: Scale-Attention Networks for Visual Correspondence
- Monocular Road Planar Parallax Estimation
- Trust, but Verify: Cross-Modality Fusion for HD Map Change Detection
- Unsupervised Abnormality Detection Using Heterogeneous Autonomous Systems
- Color Agnostic Cross-Spectral Disparity Estimation
- RayNet: Learning Volumetric 3D Reconstruction with Ray Potentials
- SRH-Net: Stacked Recurrent Hourglass Network for Stereo Matching
- StereoGAN: Bridging Synthetic-to-Real Domain Gap by Joint Optimization of Domain Translation and Stereo Matching
- PatchBatch: a Batch Augmented Loss for Optical Flow
- Deep Eyes: Binocular Depth-from-Focus on Focal Stack Pairs
- High-Performance and Tunable Stereo Reconstruction
- PointTrackNet: An End-to-End Network For 3-D Object Detection and Tracking From Point Clouds
- EDNet: Efficient Disparity Estimation with Cost Volume Combination and Attention-based Spatial Residual
- See far with TPNET: a Tile Processor and a CNN Symbiosis
- A Decomposition Model for Stereo Matching
- Expanding Sparse Guidance for Stereo Matching
- Noise-Sampling Cross Entropy Loss: Improving Disparity Regression Via Cost Volume Aware Regularizer
- MessyTable: Instance Association in Multiple Camera Views
- Robust Lane Marking Detection Algorithm Using Drivable Area Segmentation and Extended SLT
- A-TVSNet: Aggregated Two-View Stereo Network for Multi-View Stereo Depth Estimation
- Learning Image Matching by Simply Watching Video
- MSDC-Net: Multi-Scale Dense and Contextual Networks for Automated Disparity Map for Stereo Matching
- Simulating CRF with CNN for CNN
- End-to-end Learning of Cost-Volume Aggregation for Real-time Dense Stereo
- Low-level Vision by Consensus in a Spatial Hierarchy of Regions
- PlantStereo: A Stereo Matching Benchmark for Plant Surface Dense Reconstruction
- Semi-synthesis: A fast way to produce effective datasets for stereo matching
- Handcrafted and Deep Trackers: Recent Visual Object Tracking Approaches and Trends
- Learning to Localize Through Compressed Binary Maps
- Disparity Estimation of Planar Reflective Surfaces Using Specular Reflections From a Single Light Source
- Prototypicality effects in global semantic description of objects
- Robust Depth Estimation from Auto Bracketed Images
- Multi range Real-time depth inference from a monocular stabilized footage using a Fully Convolutional Neural Network
- Movement-induced Priors for Deep Stereo