10 papers
Surface Keypoint Representation for Multi-Object and Articulated Human-Object Interaction Generation
Xiaogang Peng, Zeyu Han, Zichong Meng +4
Daily activities require humans to coordinate whole-body motion with the motion of surrounding objects. Despite recent progress in human-object interaction (HOI) generation, most e…
Streaming Video Generation with Streaming Force Control
Hanhui Wang, Yiming Xie, Haiwen Feng +3
We introduce StreamForce, a streaming video generation framework that enables physically grounded control through continuous force inputs. Unlike prior video models that train sepa…
LASER: Layer-wise Scale Alignment for Training-Free Streaming 4D Reconstruction
Tianye Ding, Yiming Xie, Yiqing Liang +3
Recent feed-forward reconstruction models like VGGT and achieve impressive reconstruction quality but cannot process streaming videos due to quadratic memory complexity, lim…
Struct2D: A Perception-Guided Framework for Spatial Reasoning in MLLMs
Fangrui Zhu, Hanhui Wang, Yiming Xie +4
Unlocking spatial reasoning in Multimodal Large Language Models (MLLMs) is crucial for enabling intelligent interaction with 3D environments. While prior efforts often rely on expl…
Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked Autoregression
Zichong Meng, Yiming Xie, Xiaogang Peng +2
Since 2023, Vector Quantization (VQ)-based discrete generation methods have rapidly dominated human motion generation, primarily surpassing diffusion-based continuous generation me…
HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models
Xiaogang Peng, Yiming Xie, Zizhao Wu +3
We address the problem of generating realistic 3D human-object interactions (HOIs) driven by textual prompts. To this end, we take a modular design and decompose the complex task i…