8 papers
Structured 4D Latent Predictive Model for Robot Planning
Zhiyi Li, Peilin Wu, Xiaoshen Han +2
Video predictive models are emerging as a powerful paradigm in robotics, offering a promising path toward task generalization, long-horizon planning, and flexible decision-making.…
SceneAligner: 3D-Grounded Floorplan Localization in the Wild
Junhyeong Cho, Ruojin Cai, Hadar Averbuch-Elor
Many public buildings provide floorplans with a "you are here" indicator to help visitors orient themselves. Floorplan localization seeks to computationally replicate this capabili…
Long-tail Internet photo reconstruction
Yuan Li, Yuanbo Xiangli, Hadar Averbuch-Elor +2
Internet photo collections exhibit an extremely long-tailed distribution: a few famous landmarks are densely photographed and easily reconstructed in 3D, while most real-world site…
ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wild
Hanyu Chen, Ruojin Cai, Steve Marschner +1
Symmetry detection is a fundamental problem in computer vision, and symmetries serve as powerful priors for downstream tasks. However, existing learning-based methods for detecting…
Emergent Extreme-View Geometry in 3D Foundation Models
Yiwen Zhang, Joseph Tung, Ruojin Cai +2
3D foundation models (3DFMs) have recently transformed 3D vision, enabling joint prediction of depths, poses, and point maps directly from images. Yet their ability to reason under…
Doppelgangers++: Improved Visual Disambiguation with Geometric 3D Features
Yuanbo Xiangli, Ruojin Cai, Hanyu Chen +2
Accurate 3D reconstruction is frequently hindered by visual aliasing, where visually similar but distinct surfaces (aka, doppelgangers), are incorrectly matched. These spurious mat…