collaborators

5 papers

cs.CV2026

CRISP: Contact-Guided Real2Sim from Monocular Video with Planar Scene Primitives

Zihan Wang, Jiashun Wang, Jeff Tan +4

We introduce CRISP, a method that recovers simulatable human motion and scene geometry from monocular video. Prior work on joint human-scene reconstruction relies on data-driven pr…

cs.CV2026

MonoFusion: Sparse-View 4D Reconstruction via Monocular Fusion

Zihan Wang, Jeff Tan, Tarasha Khurana +2

We address the problem of dynamic scene reconstruction from sparse-view videos. Prior work often requires dense multi-view captures with hundreds of calibrated cameras (e.g. Panopt…

cs.RO2025

HDMI: Learning Interactive Humanoid Whole-Body Control from Human Videos

Haoyang Weng, Yitang Li, Nikhil Sobanbabu +5

Enabling robust whole-body humanoid-object interaction (HOI) remains challenging due to motion data scarcity and the contact-rich nature. We present HDMI (HumanoiD iMitation for In…

cs.CV2025

InstructPart: Task-Oriented Part Segmentation with Instruction Reasoning

Zifu Wan, Yaqi Xie, Ce Zhang +5

Large multimodal foundation models, particularly in the domains of language and vision, have significantly advanced various tasks, including robotics, autonomous driving, informati…

cs.LG2025

Impact of Noisy Supervision in Foundation Model Learning

Hao Chen, Zihan Wang, Ran Tao +5

Foundation models are usually pre-trained on large-scale datasets and then adapted to downstream tasks through tuning. However, the large-scale pre-training datasets, often inacces…