collaborators

11 papers

cs.RO2026

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control

Weili Zeng, Yitong Xing, Fulong Liu +10

The paper introduces Enfold, a method that folds the computation of a world-generative model into a predictive representation derived from the current visual scene and language ins…

cs.CV2026

Towards Consistent Video Geometry Estimation

Zhu Yu, Jingnan Gao, Runmin Zhang +9

ViGeo is a transformer-based model that estimates dense, temporally consistent geometry (depth, surface normals, and point maps) from video sequences using dynamic chunking attenti…

cs.RO2026

R3DP: Real-Time 3D-Aware Policy for Embodied Manipulation

Yuhao Zhang, Wanxi Dong, Yue Shi +13

Embodied manipulation requires accurate 3D understanding of objects and their spatial relations to plan and execute contact-rich actions. While large-scale 3D vision models provide…

cs.CV2026

SurfSplat: Conquering Feedforward 2D Gaussian Splatting with Surface Continuity Priors

Bing He, Jingnan Gao, Yunuo Chen +5

Reconstructing 3D scenes from sparse images remains a challenging task due to the difficulty of recovering accurate geometry and texture without optimization. Recent approaches lev…

cs.CV2025

POLAR: A Portrait OLAT Dataset and Generative Framework for Illumination-Aware Face Modeling

Zhuo Chen, Chengqun Yang, Zhuo Su +5

Face relighting aims to synthesize realistic portraits under novel illumination while preserving identity and geometry. However, progress remains constrained by the limited availab…

cs.CV2025

MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts

Jingnan Gao, Zhe Wang, Xianze Fang +7

Recent advances in language and vision have demonstrated that scaling up model capacity consistently improves performance across diverse tasks. In 3D visual geometry reconstruction…