Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning
arXiv:2108.00262 · doi:10.1145/3474085.3475223
Abstract
We present a generative adversarial network to synthesize 3D pose sequences of co-speech upper-body gestures with appropriate affective expressions. Our network consists of two components: a generator to synthesize gestures from a joint embedding space of features encoded from the input speech and the seed poses, and a discriminator to distinguish between the synthesized pose sequences and real 3D pose sequences. We leverage the Mel-frequency cepstral coefficients and the text transcript computed from the input speech in separate encoders in our generator to learn the desired sentiments and the associated affective cues. We design an affective encoder using multi-scale spatial-temporal graph convolutions to transform 3D pose sequences into latent, pose-based affective features. We use our affective encoder in both our generator, where it learns affective features from the seed poses to guide the gesture synthesis, and our discriminator, where it enforces the synthesized gestures to contain the appropriate affective expressions. We perform extensive evaluations on two benchmark datasets for gesture synthesis from the speech, the TED Gesture Dataset and the GENEA Challenge 2020 Dataset. Compared to the best baselines, we improve the mean absolute joint error by 10--33%, the mean acceleration difference by 8--58%, and the Fréchet Gesture Distance by 21--34%. We also conduct a user study and observe that compared to the best current baselines, around 15.28% of participants indicated our synthesized gestures appear more plausible, and around 16.32% of participants felt the gestures had more appropriate affective expressions aligned with the speech.
11 pages, 4 figures, 2 tables. Proceedings of the 29th ACM International Conference on Multimedia, October 20-24, 2021, Virtual Event, China
References in corpus (11)
- Adam: A Method for Stochastic Optimization
- Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity
- Gesticulator: A framework for semantically-aware speech-driven gesture generation
- Analyzing Input and Output Representations for Speech-Driven Gesture Generation
- Text2Gestures: A Transformer-Based Network for Generating Emotive Body Gestures for Virtual Agents
- STEP: Spatial Temporal Graph Convolutional Networks for Emotion Perception from Gaits
- A large, crowdsourced evaluation of gesture generation systems on common data: The GENEA Challenge 2020
- Speech-driven Animation with Meaningful Behaviors
- Identifying Emotions from Walking using Affective and Deep Features
- Take an Emotion Walk: Perceiving Emotions from Gaits Using Hierarchical Attention Pooling and Affective Mapping
- Generating Emotive Gaits for Virtual Agents Using Affect-Based Autoregression
Cited by in corpus (10)
- Rhythmic Gesticulator: Rhythm-Aware Co-Speech Gesture Synthesis with Hierarchical Neural Embeddings
- A Comprehensive Review of Data-Driven Co-Speech Gesture Generation
- The GENEA Challenge 2022: A large evaluation of data-driven co-speech gesture generation
- Semantic Gesticulator: Semantics-Aware Co-Speech Gesture Synthesis
- Evaluating gesture generation in a large-scale open challenge: The GENEA Challenge 2022
- BodyFormer: Semantics-guided 3D Body Gesture Synthesis with Transformer
- The Importance of Multimodal Emotion Conditioning and Affect Consistency for Embodied Conversational Agents
- DanceAnyWay: Synthesizing Beat-Guided 3D Dances with Randomized Temporal Contrastive Learning
- Speech2UnifiedExpressions: Synchronous Synthesis of Co-Speech Affective Face and Body Expressions from Affordable Inputs
- Learning Co-Speech Gesture Representations in Dialogue through Contrastive Learning: An Intrinsic Evaluation