activity
20242026
collaborators

6 papers

cs.CV2026

MotionGPT-2: A General-Purpose Motion-Language Model for Motion Generation and Understanding

Yuan Wang, Di Huang, Yaqi Zhang +5

Generating lifelike human motions from descriptive texts has experienced remarkable research focus in the recent years, propelled by the emerging requirements of digital humans.Des…

cs.CV2025

EMMA: End-to-End Multimodal Model for Autonomous Driving

Jyh-Jing Hwang, Runsheng Xu, Hubert Lin +11

We introduce EMMA, an End-to-end Multimodal Model for Autonomous driving. Built upon a multi-modal large language model foundation like Gemini, EMMA directly maps raw camera sensor…

cs.CV2025

PonderV2: Pave the Way for 3D Foundation Model with A Universal Pre-training Paradigm

Haoyi Zhu, Honghui Yang, Xiaoyang Wu +8

In contrast to numerous NLP and 2D vision foundational models, learning a 3D foundational model poses considerably greater challenges. This is primarily due to the inherent data va…

cs.CV2025

ND-SDF: Learning Normal Deflection Fields for High-Fidelity Indoor Reconstruction

Ziyu Tang, Weicai Ye, Yifan Wang +4

Neural implicit reconstruction via volume rendering has demonstrated its effectiveness in recovering dense 3D surfaces. However, it is non-trivial to simultaneously recover meticul…

cs.CV2025

Depth Any Video with Scalable Synthetic Data

Honghui Yang, Di Huang, Wei Yin +6

Video depth estimation has long been hindered by the scarcity of consistent and scalable ground truth data, leading to inconsistent and unreliable results. In this paper, we introd…

cs.CV2024

NeuRodin: A Two-stage Framework for High-Fidelity Neural Surface Reconstruction

Yifan Wang, Di Huang, Weicai Ye +3

Signed Distance Function (SDF)-based volume rendering has demonstrated significant capabilities in surface reconstruction. Although promising, SDF-based methods often fail to captu…