10 papers
StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring
Xiangyue Zhang, Jianfang Li, Jiaxu Zhang +2
Real-time co-speech gesture generation must produce 3D motion clip by clip as speech arrives. Existing streaming methods are open-loop: each clip depends on past context, but the m…
DC-Motion: Decoupling Structure and Details via Discrete-Continuous Tokens for Human Motion Generation
Hequan Wang, Xuean Chen, Jiaxu Zhang +2
Text-to-motion generation requires modeling both global action structure and fine-grained motion dynamics from natural language. Existing approaches typically rely on either contin…
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
Shihao Cheng, Jiaxu Zhang, Quanyue Song +6
Motion, speech, and sound effects are fundamental elements of human-centric videos, yet their heterogeneous temporal characteristics make joint generation highly challenging. Exist…
OmniSVG: A Unified Scalable Vector Graphics Generation Model
Yiying Yang, Wei Cheng, Sijin Chen +7
Scalable Vector Graphics (SVG) is an important image format widely adopted in graphic design because of their resolution independence and editability. The study of generating high-…
DreamDance: Animating Character Art via Inpainting Stable Gaussian Worlds
Jiaxu Zhang, Xianfang Zeng, Xin Chen +4
This paper presents DreamDance, a novel character art animation framework capable of producing stable, consistent character and scene motion conditioned on precise camera trajector…
EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation
Xiangyue Zhang, Jianfang Li, Jiaxu Zhang +3
Masked modeling has shown promise in co-speech gesture generation. However, it struggles to identify semantically significant frames for effective motion masking. In this work, we…