most citedMitigating Error Accumulation in Co-Speech Motion Generation via Global Rotation Diffusion and Multi-Level Constraints

1 citations · 1 across the 5 of their papers we have counts for

collaborators

7 papers

cs.CV2026

StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring

Xiangyue Zhang, Jianfang Li, Jiaxu Zhang +2

Real-time co-speech gesture generation must produce 3D motion clip by clip as speech arrives. Existing streaming methods are open-loop: each clip depends on past context, but the m…

eess.AS2026

The WER Trap: Shattering the Illusion of Unified Tokens in Speech Language Models

Xiangyu Zhang, Yuxin Li, Haoyang Zhang +5

The pursuit of a "unified" discrete token for both speech understanding and generation has led the Speech Language Model (SLM) community to heavily rely on Word Error Rate (WER) --…

cs.CV2026

PersonaGesture: Single-Reference Co-Speech Gesture Personalization for Unseen Speakers

Xiangyue Zhang, Yiyi Cai, Kunhang Li +6

We propose PersonaGesture, a diffusion-based pipeline for single-reference co-speech gesture personalization of unseen speakers. Given target speech and one motion clip from a new…

cs.CV2026

Not All Frames Are Equal: Complexity-Aware Masked Motion Generation via Motion Spectral Descriptors

Pengfei Zhou, Xiangyue Zhang, Xukun Shen +1

Masked generative models have become a strong paradigm for text-to-motion synthesis, but they still treat motion frames too uniformly during masking, attention, and decoding. This…

cs.CV20261 cited

Mitigating Error Accumulation in Co-Speech Motion Generation via Global Rotation Diffusion and Multi-Level Constraints

Xiangyue Zhang, Jianfang Li, Jianqiang Ren +1

Reliable long-horizon co-speech gesture generation requires precise motion representation and consistent structural priors across all joints. Existing generative methods typically…

cs.GR2025

EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation

Xiangyue Zhang, Jianfang Li, Jiaxu Zhang +3

Masked modeling has shown promise in co-speech gesture generation. However, it struggles to identify semantically significant frames for effective motion masking. In this work, we…