Rhythmic Gesticulator: Rhythm-Aware Co-Speech Gesture Synthesis with Hierarchical Neural Embeddings
arXiv:2210.01448 · doi:10.1145/3550454.3555435
Abstract
Automatic synthesis of realistic co-speech gestures is an increasingly important yet challenging task in artificial embodied agent creation. Previous systems mainly focus on generating gestures in an end-to-end manner, which leads to difficulties in mining the clear rhythm and semantics due to the complex yet subtle harmony between speech and gestures. We present a novel co-speech gesture synthesis method that achieves convincing results both on the rhythm and semantics. For the rhythm, our system contains a robust rhythm-based segmentation pipeline to ensure the temporal coherence between the vocalization and gestures explicitly. For the gesture semantics, we devise a mechanism to effectively disentangle both low- and high-level neural embeddings of speech and motion based on linguistic theory. The high-level embedding corresponds to semantics, while the low-level embedding relates to subtle variations. Lastly, we build correspondence between the hierarchical embeddings of the speech and the motion, resulting in rhythm- and semantics-aware gesture synthesis. Evaluations with existing objective metrics, a newly proposed rhythmic metric, and human feedback show that our method outperforms state-of-the-art systems by a clear margin.
SIGGRAPH Asia 2022 (Journal Track); Project Page: https://pku-mocca.github.io/Rhythmic-Gesticulator-Page/
References in corpus (7)
- Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity
- vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
- Character Controllers Using Motion VAEs
- Robust Motion In-betweening
- Analyzing Input and Output Representations for Speech-Driven Gesture Generation
- A large, crowdsourced evaluation of gesture generation systems on common data: The GENEA Challenge 2020
- Speech2Properties2Gestures: Gesture-Property Prediction as a Tool for Generating Representational Gestures from Speech
Cited by in corpus (14)
- Listen, Denoise, Action! Audio-Driven Motion Synthesis with Diffusion Models
- InterGen: Diffusion-based Multi-human Motion Generation under Complex Interactions
- A Comprehensive Review of Data-Driven Co-Speech Gesture Generation
- Semantic Gesticulator: Semantics-Aware Co-Speech Gesture Synthesis
- Evaluating gesture generation in a large-scale open challenge: The GENEA Challenge 2022
- BodyFormer: Semantics-guided 3D Body Gesture Synthesis with Transformer
- UnifiedGesture: A Unified Gesture Synthesis Model for Multiple Skeletons
- SATO: Stable Text-to-Motion Framework
- GesGPT: Speech Gesture Synthesis With Text Parsing from ChatGPT
- EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation
- ExpGest: Expressive Speaker Generation Using Diffusion Model and Hybrid Audio-Text Guidance
- Speech2UnifiedExpressions: Synchronous Synthesis of Co-Speech Affective Face and Body Expressions from Affordable Inputs
- Semantics-Aware Human Motion Generation from Audio Instructions
- Script2Screen: Supporting Dialogue Scriptwriting with Interactive Audiovisual Generation