Transformer Network for Semantically-Aware and Speech-Driven Upper-Face Generation
arXiv:2110.04527
Abstract
We propose a semantically-aware speech driven model to generate expressive and natural upper-facial and head motion for Embodied Conversational Agents (ECA). In this work, we aim to produce natural and continuous head motion and upper-facial gestures synchronized with speech. We propose a model that generates these gestures based on multimodal input features: the first modality is text, and the second one is speech prosody. Our model makes use of Transformers and Convolutions to map the multimodal features that correspond to an utterance to continuous eyebrows and head gestures. We conduct subjective and objective evaluations to validate our approach and compare it with state of the art.
References in corpus (6)
- A Survey on Methods and Theories of Quantized Neural Networks
- Transformers with convolutional context for ASR
- A Review of Evaluation Practices of Gesture Generation in Embodied Conversational Agents
- Speech-driven Animation with Meaningful Behaviors
- Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion
- Prediction of head motion from speech waveforms with a canonical-correlation-constrained autoencoder