4 papers
SmartMage: Dynamic Modality Orchestration for 3D Scene Understanding
Yue Zhang, Yingzhao Jian, Yunqiu Xu +2
Understanding 3D scenes is fundamental to embodied intelligence, requiring joint reasoning over heterogeneous information from multiple modalities, including visual and geometric c…
Depth Estimators Are Implicit Neural Fields for 3D Scene Geometry Inpainting and Reconstruction
Yingzhao Jian, Zihao Lin, Hehe Fan
The 3D geometry of real-world scene data is often incomplete. Mainstream methods use depth estimators to inpaint missing structure. However, their prediction results can be inconsi…
Endowing GPT-4 with a Humanoid Body: Building the Bridge Between Off-the-Shelf VLMs and the Physical World
Yingzhao Jian, Zhongan Wang, Yi Yang +1
Humanoid agents often struggle to handle flexible and diverse interactions in open environments. A common solution is to collect massive datasets to train a highly capable model, b…
Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts
Yue Zhang, Yingzhao Jian, Hehe Fan +2
Recent advancements in multimodal large language models (MLLMs) have demonstrated considerable potential for comprehensive 3D scene understanding. However, existing approaches typi…