1 citations · 1 across the 5 of their papers we have counts for
7 papers
StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring
Xiangyue Zhang, Jianfang Li, Jiaxu Zhang +2
Real-time co-speech gesture generation must produce 3D motion clip by clip as speech arrives. Existing streaming methods are open-loop: each clip depends on past context, but the m…
The WER Trap: Shattering the Illusion of Unified Tokens in Speech Language Models
Xiangyu Zhang, Yuxin Li, Haoyang Zhang +5
The pursuit of a "unified" discrete token for both speech understanding and generation has led the Speech Language Model (SLM) community to heavily rely on Word Error Rate (WER) --…
PersonaGesture: Single-Reference Co-Speech Gesture Personalization for Unseen Speakers
Xiangyue Zhang, Yiyi Cai, Kunhang Li +6
We propose PersonaGesture, a diffusion-based pipeline for single-reference co-speech gesture personalization of unseen speakers. Given target speech and one motion clip from a new…
Not All Frames Are Equal: Complexity-Aware Masked Motion Generation via Motion Spectral Descriptors
Pengfei Zhou, Xiangyue Zhang, Xukun Shen +1
Masked generative models have become a strong paradigm for text-to-motion synthesis, but they still treat motion frames too uniformly during masking, attention, and decoding. This…
Mitigating Error Accumulation in Co-Speech Motion Generation via Global Rotation Diffusion and Multi-Level Constraints
Xiangyue Zhang, Jianfang Li, Jianqiang Ren +1
Reliable long-horizon co-speech gesture generation requires precise motion representation and consistent structural priors across all joints. Existing generative methods typically…
EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation
Xiangyue Zhang, Jianfang Li, Jiaxu Zhang +3
Masked modeling has shown promise in co-speech gesture generation. However, it struggles to identify semantically significant frames for effective motion masking. In this work, we…