10 papers
Honey, I Shrunk the Arc de Triomphe!
Yuanbo Xiangli, Hanyu Chen, Xueqing Tsang +1
Metric scale monocular geometry estimation has seen significant progress through large-scale data aggregation, yet current foundation models suffer from a persistent ''scale-collap…
CityRAG: Stepping Into a City via Spatially-Grounded Video Generation
Gene Chou, Charles Herrmann, Kyle Genova +6
We address the problem of generating a 3D-consistent, navigable environment that is spatially grounded: a simulation of a real location. Existing video generative models can produc…
G3T Up! Gravity Aligned Coordinate Frames Simplify Pointmap Processing
Bharath Raj Nagoor Kani, Noah Snavely
Modern feed-forward 3D reconstruction methods like VGGT predict pixel-aligned pointmaps in camera-centric coordinate frames. However, this choice of coordinate frame is not always…
Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly
Aditya Chetan, Eric Cai, Peeyush Kushwaha +5
The emergence of Large Vision-Language Models (LVLMs) has significantly advanced video understanding capabilities. However, existing benchmarks focus predominantly on coarse-graine…
Long-tail Internet photo reconstruction
Yuan Li, Yuanbo Xiangli, Hadar Averbuch-Elor +2
Internet photo collections exhibit an extremely long-tailed distribution: a few famous landmarks are densely photographed and easily reconstructed in 3D, while most real-world site…
ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wild
Hanyu Chen, Ruojin Cai, Steve Marschner +1
Symmetry detection is a fundamental problem in computer vision, and symmetries serve as powerful priors for downstream tasks. However, existing learning-based methods for detecting…