Text2Gestures: A Transformer-Based Network for Generating Emotive Body Gestures for Virtual Agents
arXiv:2101.11101 · doi:10.1109/VR50410.2021.00037
Abstract
We present Text2Gestures, a transformer-based learning method to interactively generate emotive full-body gestures for virtual agents aligned with natural language text inputs. Our method generates emotionally expressive gestures by utilizing the relevant biomechanical features for body expressions, also known as affective features. We also consider the intended task corresponding to the text and the target virtual agents' intended gender and handedness in our generation pipeline. We train and evaluate our network on the MPI Emotional Body Expressions Database and observe that our network produces state-of-the-art performance in generating gestures for virtual agents aligned with the text for narration or conversation. Our network can generate these gestures at interactive rates on a commodity GPU. We conduct a web-based user study and observe that around 91% of participants indicated our generated gestures to be at least plausible on a five-point Likert Scale. The emotions perceived by the participants from the gestures are also strongly positively correlated with the corresponding intended emotions, with a minimum Pearson coefficient of 0.77 in the valence dimension.
10 pages, 8 figures, 2 tables
References in corpus (7)
- Speech Gesture Generation from the Trimodal Context of Text, Audio, and Speaker Identity
- Gesticulator: A framework for semantically-aware speech-driven gesture generation
- Analyzing Input and Output Representations for Speech-Driven Gesture Generation
- STEP: Spatial Temporal Graph Convolutional Networks for Emotion Perception from Gaits
- Speech-driven Animation with Meaningful Behaviors
- Take an Emotion Walk: Perceiving Emotions from Gaits Using Hierarchical Attention Pooling and Affective Mapping
- Generating Emotive Gaits for Virtual Agents Using Affect-Based Autoregression
Cited by in corpus (6)
- Speech2AffectiveGestures: Synthesizing Co-Speech Gestures with Generative Adversarial Affective Expression Learning
- Semantic Gesticulator: Semantics-Aware Co-Speech Gesture Synthesis
- AI Choreographer: Music Conditioned 3D Dance Generation with AIST++
- The Importance of Multimodal Emotion Conditioning and Affect Consistency for Embodied Conversational Agents
- fMRI2GES: Co-speech Gesture Reconstruction from fMRI Signal with Dual Brain Decoding Alignment
- Speech2UnifiedExpressions: Synchronous Synthesis of Co-Speech Affective Face and Body Expressions from Affordable Inputs