From the 1 of 7 linked papers with an AI index.
7 papers
Towards Consistent Video Geometry Estimation
Zhu Yu, Jingnan Gao, Runmin Zhang +9
ViGeo is a transformer-based model that estimates dense, temporally consistent geometry (depth, surface normals, and point maps) from video sequences using dynamic chunking attenti…
Glob3R: Global Structure-from-Motion with 3D Foundation Models
Junyuan Deng, Heng Li, Kejie Qiu +7
Recent 3D geometric foundation models, such as VGGT, provide robust feed-forward 3D reconstruction by directly predicting camera poses and 3D scene points from input images. Howeve…
Large Depth Completion Model from Sparse Observations
Zhu Yu, Zhengyi Zhao, Runmin Zhang +7
This work presents the Large Depth Completion Model (LDCM), a simple, effective, and robust framework for single-view metric depth estimation with sparse observations. Without rely…
The Thinking Pixel: Recursive Sparse Reasoning in Multimodal Diffusion Latents
Yuwei Sun, Yuxuan Yao, Hui Li +1
Diffusion models have achieved success in high-fidelity data synthesis, yet their capacity for more complex, structured reasoning like text following tasks remains constrained. Whi…
LHM++: An Efficient Large Human Reconstruction Model for Pose-free Images to 3D
Lingteng Qiu, Peihao Li, Heyuan Li +9
Reconstructing animatable 3D humans from casually captured images of articulated subjects without camera or pose information is highly practical but remains challenging due to view…
OmniMotion: Multimodal Motion Generation with Continuous Masked Autoregression
Zhe Li, Weihao Yuan, Weichao Shen +3
Whole-body multi-modal human motion generation poses two primary challenges: creating an effective motion generation mechanism and integrating various modalities, such as text, spe…