5 papers
EVA01: Unified Native 3D Understanding and Generation via Mixture-of-Transformers
Zongyuan Yang, Mingjing Yi, Wanli Ma +8
This paper addresses the challenge of integrating 3D meshes as a native modality within Multimodal Large Language Models (MLLMs). Diffusion-based large reconstruction models decoup…
EmbodMocap: In-the-Wild 4D Human-Scene Reconstruction for Embodied Agents
Wenjia Wang, Liang Pan, Huaijin Pi +8
Human behaviors in the real world naturally encode rich, long-term contextual information that can be leveraged to train embodied agents for perception, understanding, and acting.…
CBIL: Collective Behavior Imitation Learning for Fish from Real Videos
Yifan Wu, Zhiyang Dou, Yuko Ishiwaka +5
Reproducing realistic collective behaviors presents a captivating yet formidable challenge. Traditional rule-based methods rely on hand-crafted principles, limiting motion diversit…
Zero-Shot Human-Object Interaction Synthesis with Multimodal Priors
Yuke Lou, Yiming Wang, Zhen Wu +4
Human-object interaction (HOI) synthesis is important for various applications, ranging from virtual reality to robotics. However, acquiring 3D HOI data is challenging due to its c…
SIMS: Simulating Stylized Human-Scene Interactions with Retrieval-Augmented Script Generation
Wenjia Wang, Liang Pan, Zhiyang Dou +7
Simulating stylized human-scene interactions (HSI) in physical environments is a challenging yet fascinating task. Prior works emphasize long-term execution but fall short in achie…