From the 1 of 8 linked papers with an AI index.
8 papers
GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure
Leslie Gu, Junhwa Hur, Charles Herrmann +4
GeCo is a geometry-based metric that detects deformation and occlusion inconsistencies in generated videos by combining residual motion and depth cues, providing dense consistency…
RiGS: Rigid-aware 4D Gaussian Splatting from a Single Monocular Video
Chenyu Wu, Wanhua Li, Zhu-Tian Chen +1
Reconstructing dynamic 3D scenes from monocular videos is a fundamental yet highly challenging task, as real-world motions often involve both long-term smooth transformations and s…
LangFlash: Feed-forward 3D Language Gaussian Splatting from Sparse Unposed Images
Yilong Liu, Wanhua Li, Chen Zhu-Tian +1
We present LangFlash, a feed-forward framework for 3D Language Gaussian Splatting that reconstructs 3D scenes parameterized by Gaussian primitives enriched with language-aligned se…
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
Yifan Liu, Fangneng Zhan, Kaichen Zhou +3
Vision-language models (VLMs) struggle with 3D-related tasks such as spatial cognition and physical understanding, which are crucial for real-world applications like robotics and e…
Advances in Feed-Forward 3D Reconstruction and View Synthesis: A Survey
Jiahui Zhang, Yuelei Li, Anpei Chen +15
3D reconstruction and view synthesis are foundational problems in computer vision, graphics, and immersive technologies such as augmented reality (AR), virtual reality (VR), and di…
AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
Tianling Xu, Shengzhe Gan, Leslie Gu +3
Active 3D reconstruction enables an agent to autonomously select viewpoints to efficiently obtain accurate and complete scene geometry, rather than passively reconstructing scenes…