1 citations · 3 across the 10 of their papers we have counts for
4 papers · 1 filter
Glob3R: Global Structure-from-Motion with 3D Foundation Models
Junyuan Deng, Heng Li, Kejie Qiu +7
Recent 3D geometric foundation models, such as VGGT, provide robust feed-forward 3D reconstruction by directly predicting camera poses and 3D scene points from input images. Howeve…
Large Depth Completion Model from Sparse Observations
Zhu Yu, Zhengyi Zhao, Runmin Zhang +7
This work presents the Large Depth Completion Model (LDCM), a simple, effective, and robust framework for single-view metric depth estimation with sparse observations. Without rely…
Towards Consistent Video Geometry Estimation
Zhu Yu, Jingnan Gao, Runmin Zhang +9
This work presents ViGeo, a feed-forward foundation model for recovering spatially dense and temporally consistent geometry from video sequences. Built upon a plain transformer arc…
The Thinking Pixel: Recursive Sparse Reasoning in Multimodal Diffusion Latents
Yuwei Sun, Yuxuan Yao, Hui Li +1
Diffusion models have achieved success in high-fidelity data synthesis, yet their capacity for more complex, structured reasoning like text following tasks remains constrained. Whi…