Emotional Speech-Driven Animation with Content-Emotion Disentanglement
arXiv:2306.08990 · doi:10.1145/3610548.3618183
Abstract
To be widely adopted, 3D facial avatars must be animated easily, realistically, and directly from speech signals. While the best recent methods generate 3D animations that are synchronized with the input audio, they largely ignore the impact of emotions on facial expressions. Realistic facial animation requires lip-sync together with the natural expression of emotion. To that end, we propose EMOTE (Expressive Model Optimized for Talking with Emotion), which generates 3D talking-head avatars that maintain lip-sync from speech while enabling explicit control over the expression of emotion. To achieve this, we supervise EMOTE with decoupled losses for speech (i.e., lip-sync) and emotion. These losses are based on two key observations: (1) deformations of the face due to speech are spatially localized around the mouth and have high temporal frequency, whereas (2) facial expressions may deform the whole face and occur over longer intervals. Thus, we train EMOTE with a per-frame lip-reading loss to preserve the speech-dependent content, while supervising emotion at the sequence level. Furthermore, we employ a content-emotion exchange mechanism in order to supervise different emotions on the same audio, while maintaining the lip motion synchronized with the speech. To employ deep perceptual losses without getting undesirable artifacts, we devise a motion prior in the form of a temporal VAE. Due to the absence of high-quality aligned emotional 3D face datasets with speech, EMOTE is trained with 3D pseudo-ground-truth extracted from an emotional video dataset (i.e., MEAD). Extensive qualitative and perceptual evaluations demonstrate that EMOTE produces speech-driven facial animations with better lip-sync than state-of-the-art methods trained on the same data, while offering additional, high-quality emotional control.
SIGGRAPH Asia 2023 Conference Paper
References in corpus (4)
Cited by in corpus (12)
- EmoFace: Audio-driven Emotional 3D Face Animation
- ProbTalk3D: Non-Deterministic Emotion Controllable Speech-Driven 3D Facial Animation Synthesis Using VQ-VAE
- Evaluation of Generative Models for Emotional 3D Animation Generation in VR
- EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation
- A Platform for Interactive AI Character Experiences
- FreeAvatar: Robust 3D Facial Animation Transfer by Learning an Expression Foundation Model
- Model See Model Do: Speech-Driven Facial Animation with Style Control
- InterAct: A Large-Scale Dataset of Dynamic, Expressive and Interactive Activities between Two People in Daily Scenarios
- Instruction-Driven 3D Facial Expression Generation and Transition
- Tiny is not small enough: High-quality, low-resource facial animation models through hybrid knowledge distillation
- Advancing Talking Head Generation: A Comprehensive Survey of Multi-Modal Methodologies, Datasets, Evaluation Metrics, and Loss Functions
- Spline-based Transformers