11 papers
Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control
Weili Zeng, Yitong Xing, Fulong Liu +10
The paper introduces Enfold, a method that folds the computation of a world-generative model into a predictive representation derived from the current visual scene and language ins…
Towards Consistent Video Geometry Estimation
Zhu Yu, Jingnan Gao, Runmin Zhang +9
ViGeo is a transformer-based model that estimates dense, temporally consistent geometry (depth, surface normals, and point maps) from video sequences using dynamic chunking attenti…
R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation
Yuhao Zhang, Wanxi Dong, Yue Shi +13
Embodied manipulation requires accurate 3D understanding of objects and their spatial relations to plan and execute contact-rich actions. While large-scale 3D vision models provide…
SurfSplat: Conquering Feedforward 2D Gaussian Splatting with Surface Continuity Priors
Bing He, Jingnan Gao, Yunuo Chen +5
Reconstructing 3D scenes from sparse images remains a challenging task due to the difficulty of recovering accurate geometry and texture without optimization. Recent approaches lev…
POLAR: A Portrait OLAT Dataset and Generative Framework for Illumination-Aware Face Modeling
Zhuo Chen, Chengqun Yang, Zhuo Su +5
Face relighting aims to synthesize realistic portraits under novel illumination while preserving identity and geometry. However, progress remains constrained by the limited availab…
MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts
Jingnan Gao, Zhe Wang, Xianze Fang +7
Recent advances in language and vision have demonstrated that scaling up model capacity consistently improves performance across diverse tasks. In 3D visual geometry reconstruction…