10 papers
GemDepth: Geometry-Embedded Features for 3D-Consistent Video Depth
Yuecheng Liu, Junda Cheng, Longliang Liu +4
Video depth estimation extends monocular prediction into the temporal domain to ensure coherence. However, existing methods often suffer from spatial blurring in fine-detail region…
PointForward: Feedforward Driving Reconstruction through Point-Aligned Representations
Cheng Chi, Xianqi Wang, Hongcheng Luo +9
High-fidelity reconstruction of driving scenes is crucial for autonomous driving. While recent feedforward 3D Gaussian Splatting (3DGS) methods enable fast reconstruction, their pe…
PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts
Xianqi Wang, Hao Yang, Hangtian Wang +4
Modern stereo matching methods have leveraged monocular depth foundation models to achieve superior zero-shot generalization performance. However, most existing methods primarily f…
Generalized Geometry Encoding Volume for Real-time Stereo Matching
Jiaxin Liu, Gangwei Xu, Xianqi Wang +2
Real-time stereo matching methods primarily focus on enhancing in-domain performance but often overlook the critical importance of generalization in real-world applications. In con…
Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers
Gangwei Xu, Haotong Lin, Hongcheng Luo +11
This paper presents Pixel-Perfect Depth, a monocular depth estimation model based on pixel-space diffusion generation that produces high-quality, flying-pixel-free point clouds fro…
MonSter++: Unified Stereo Matching, Multi-view Stereo, and Real-time Stereo with Monodepth Priors
Junda Cheng, Wenjing Liao, Zhipeng Cai +10
We introduce MonSter++, a geometric foundation model for multi-view depth estimation, unifying rectified stereo matching and unrectified multi-view stereo. Both tasks fundamentally…