12 papers
ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
Yukang Cao, Haozhe Xie, Beichen Wen +13
The paper presents ACE, a data collection system that records synchronized multimodal streams—including egocentric and multi-view video, full-body and hand motion, object geometry,…
Collaborative Multi-Modal Coding for High-Quality 3D Generation
Ziang Cao, Zhaoxi Chen, Liang Pan +1
3D content inherently encompasses multi-modal characteristics and can be projected into different modalities (e.g., RGB images, RGBD, and point clouds). Each modality exhibits dist…
PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects
Ziang Cao, Yinghao Liu, Haitian Li +5
Simulation-ready physical 3D assets have emerged as a promising direction owing to their broad applicability in downstream tasks. However, most existing 3D generation methods eithe…
Is Your Driving World Model an All-Around Player?
Lingdong Kong, Ao Liang, Tianyi Yan +20
Today's driving world models can generate remarkably realistic dash-cam videos, yet no single model excels universally. Some generate photorealistic textures but violate basic phys…
Scaling Spatial Intelligence with Multimodal Foundation Models
Zhongang Cai, Ruisi Wang, Chenyang Gu +26
Despite remarkable progress, multimodal foundation models still exhibit surprising deficiencies in spatial intelligence. In this work, we explore scaling up multimodal foundation m…
HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions
Yukang Cao, Haozhe Xie, Fangzhou Hong +4
We present HSImul3R, a unified framework for simulation-ready 3D reconstruction of human-scene interactions (HSI) from casual captures, including sparse-view images and monocular v…