9 papers
DC-Motion: Decoupling Structure and Details via Discrete-Continuous Tokens for Human Motion Generation
Hequan Wang, Xuean Chen, Jiaxu Zhang +2
Text-to-motion generation requires modeling both global action structure and fine-grained motion dynamics from natural language. Existing approaches typically rely on either contin…
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
Shihao Cheng, Jiaxu Zhang, Quanyue Song +6
Motion, speech, and sound effects are fundamental elements of human-centric videos, yet their heterogeneous temporal characteristics make joint generation highly challenging. Exist…
Uni-HOI:A Unified framework for Learning the Joint distribution of Text and Human-Object Interaction
Mengfei Zhang, Jinlu Zhang, Zhigang Tu
Modeling 4D human-object interaction (HOI) is a compelling challenge in computer vision and an essential technology powering virtual and mixed-reality applications. While existing…
DreamDance: Animating Character Art via Inpainting Stable Gaussian Worlds
Jiaxu Zhang, Xianfang Zeng, Xin Chen +4
This paper presents DreamDance, a novel character art animation framework capable of producing stable, consistent character and scene motion conditioned on precise camera trajector…
EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation
Xiangyue Zhang, Jianfang Li, Jiaxu Zhang +3
Masked modeling has shown promise in co-speech gesture generation. However, it struggles to identify semantically significant frames for effective motion masking. In this work, we…
SemTalk: Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis
Xiangyue Zhang, Jianfang Li, Jiaxu Zhang +4
Co-speech gesture generation must carefully integrate common rhythmic motion with rare yet essential semantic gestures. In this work, we propose SemTalk for holistic co-speech gest…