7 papers
RGB-Pointmap Pretraining for Unified 3D Scene Understanding
Ye Mao, Weixun Luo, Ranran Huang +2
Pretraining 3D encoders through alignment with Contrastive Language-Image Pre-training (CLIP) has emerged as a promising direction for learning generalizable representations for 3D…
From None to All: Self-Supervised 3D Reconstruction via Novel View Synthesis
Ranran Huang, Weixun Luo, Ye Mao +1
In this paper, we introduce NAS3R, a self-supervised feed-forward framework that jointly learns explicit 3D geometry and camera parameters with no ground-truth annotations and no p…
POMA-3D: The Point Map Way to 3D Scene Understanding
Ye Mao, Weixun Luo, Ranran Huang +2
In this paper, we introduce POMA-3D, the first self-supervised 3D representation model learned from point maps. Point maps encode explicit 3D coordinates on a structured 2D grid, p…
Lite Any Stereo: Efficient Zero-Shot Stereo Matching
Junpeng Jing, Weixun Luo, Ye Mao +1
Recent advances in stereo matching have focused on accuracy, often at the cost of significantly increased model size. Traditionally, the community has regarded efficient models as…
Match Stereo Videos via Bidirectional Alignment
Junpeng Jing, Ye Mao, Anlan Qiu +1
Video stereo matching is the task of estimating consistent disparity maps from rectified stereo videos. There is considerable scope for improvement in both datasets and methods wit…
Stereo Any Video: Temporally Consistent Stereo Matching
Junpeng Jing, Weixun Luo, Ye Mao +1
This paper introduces Stereo Any Video, a powerful framework for video stereo matching. It can estimate spatially accurate and temporally consistent disparities without relying on…