Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity
arXiv:2009.02119 · doi:10.1145/3414685.3417838
Abstract
For human-like agents, including virtual avatars and social robots, making proper gestures while speaking is crucial in human--agent interaction. Co-speech gestures enhance interaction experiences and make the agents look alive. However, it is difficult to generate human-like gestures due to the lack of understanding of how people gesture. Data-driven approaches attempt to learn gesticulation skills from human demonstrations, but the ambiguous and individual nature of gestures hinders learning. In this paper, we present an automatic gesture generation model that uses the multimodal context of speech text, audio, and speaker identity to reliably generate gestures. By incorporating a multimodal context and an adversarial training scheme, the proposed model outputs gestures that are human-like and that match with speech content and rhythm. We also introduce a new quantitative evaluation metric for gesture generation models. Experiments with the introduced metric and subjective human evaluation showed that the proposed gesture generation model is better than existing end-to-end generation models. We further confirm that our model is able to work with synthesized audio in a scenario where contexts are constrained, and show that different gesture styles can be generated for the same speech by specifying different speaker identities in the style embedding space that is learned from videos of various speakers. All the code and data is available at https://github.com/ai4r/Gesture-Generation-from-Trimodal-Context.
16 pages; ACM Transactions on Graphics (SIGGRAPH Asia 2020)
References in corpus (4)
Cited by in corpus (38)
- Listen, Denoise, Action! Audio-Driven Motion Synthesis with Diffusion Models
- Text2Gestures: A Transformer-Based Network for Generating Emotive Body Gestures for Virtual Agents
- Rhythmic Gesticulator: Rhythm-Aware Co-Speech Gesture Synthesis with Hierarchical Neural Embeddings
- A Comprehensive Review of Data-Driven Co-Speech Gesture Generation
- Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning
- The GENEA Challenge 2022: A large evaluation of data-driven co-speech gesture generation
- A Review of Evaluation Practices of Gesture Generation in Embodied Conversational Agents
- Driving-Signal Aware Full-Body Avatars
- A large, crowdsourced evaluation of gesture generation systems on common data: The GENEA Challenge 2020
- Semantic Gesticulator: Semantics-Aware Co-Speech Gesture Synthesis
- The Ethical Implications of Generative Audio Models: A Systematic Literature Review
- The DiffuseStyleGesture+ entry to the GENEA Challenge 2023
- Speaker Extraction with Co-Speech Gestures Cue
- Evaluating gesture generation in a large-scale open challenge: The GENEA Challenge 2022
- Speech2Properties2Gestures: Gesture-Property Prediction as a Tool for Generating Representational Gestures from Speech
- SGToolkit: An Interactive Gesture Authoring Toolkit for Embodied Conversational Agents
- BodyFormer: Semantics-guided 3D Body Gesture Synthesis with Transformer
- UnifiedGesture: A Unified Gesture Synthesis Model for Multiple Skeletons
- Learning to Generate Pointing Gestures in Situated Embodied Conversational Agents
- The ReprGesture entry to the GENEA Challenge 2022
- Integrated Speech and Gesture Synthesis
- The Importance of Multimodal Emotion Conditioning and Affect Consistency for Embodied Conversational Agents
- GesGPT: Speech Gesture Synthesis With Text Parsing from ChatGPT
- Human Motion Video Generation: A Survey
- Evaluation of Generative Models for Emotional 3D Animation Generation in VR
- EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation
- Speech-Gesture GAN: Gesture Generation for Robots and Embodied Agents
- MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture Generation
- DanceAnyWay: Synthesizing Beat-Guided 3D Dances with Randomized Temporal Contrastive Learning
- ExpGest: Expressive Speaker Generation Using Diffusion Model and Hybrid Audio-Text Guidance
- fMRI2GES: Co-speech Gesture Reconstruction from fMRI Signal with Dual Brain Decoding Alignment
- Speech2UnifiedExpressions: Synchronous Synthesis of Co-Speech Affective Face and Body Expressions from Affordable Inputs
- Speech2Video: Cross-Modal Distillation for Speech to Video Generation
- Learning Co-Speech Gesture Representations in Dialogue through Contrastive Learning: An Intrinsic Evaluation
- Multi-Resolution Generative Modeling of Human Motion from Limited Data
- TranSTYLer: Multimodal Behavioral Style Transfer for Facial and Body Gestures Generation
- Multimodal Quantitative Measures for Multiparty Behaviour Evaluation
- META4: Semantically-Aligned Generation of Metaphoric Gestures Using Self-Supervised Text and Speech Representation