13 citations · 13 across the 10 of their papers we have counts for
10 papers · 1 filter
AVSplat: Dense-View Feed-Forward 3D Gaussian Splatting with Assist-View Preconditioning
Muyu Xu, Fangneng Zhan, Yu Wei +2
Pose-free feed-forward 3D Gaussian Splatting enables novel view synthesis from uncalibrated multi-view images. Although more views should improve performance, existing methods ofte…
RiGS: Rigid-aware 4D Gaussian Splatting from a Single Monocular Video
Chenyu Wu, Wanhua Li, Zhu-Tian Chen +1
Reconstructing dynamic 3D scenes from monocular videos is a fundamental yet highly challenging task, as real-world motions often involve both long-term smooth transformations and s…
LangFlash: Feed-forward 3D Language Gaussian Splatting from Sparse Unposed Images
Yilong Liu, Wanhua Li, Chen Zhu-Tian +1
We present LangFlash, a feed-forward framework for 3D Language Gaussian Splatting that reconstructs 3D scenes parameterized by Gaussian primitives enriched with language-aligned se…
GeCo: Evaluating Geometric Consistency for Video Generation via Motion and Structure
Leslie Gu, Junhwa Hur, Charles Herrmann +4
We introduce GeCo, a geometry-grounded metric for jointly detecting geometric deformation and occlusion-inconsistency artifacts in static scenes. By fusing residual motion and dept…
AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
Tianling Xu, Shengzhe Gan, Leslie Gu +3
Active 3D reconstruction enables an agent to autonomously select viewpoints to efficiently obtain accurate and complete scene geometry, rather than passively reconstructing scenes…
Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
Yifan Liu, Fangneng Zhan, Kaichen Zhou +3
Vision-language models (VLMs) struggle with 3D-related tasks such as spatial cognition and physical understanding, which are crucial for real-world applications like robotics and e…