A Lip Sync Expert Is All You Need for Speech to Lip Generation In The Wild
arXiv:2008.10010 · doi:10.1145/3394171.3413532
Abstract
In this work, we investigate the problem of lip-syncing a talking face video of an arbitrary identity to match a target speech segment. Current works excel at producing accurate lip movements on a static image or videos of specific people seen during the training phase. However, they fail to accurately morph the lip movements of arbitrary identities in dynamic, unconstrained talking face videos, resulting in significant parts of the video being out-of-sync with the new audio. We identify key reasons pertaining to this and hence resolve them by learning from a powerful lip-sync discriminator. Next, we propose new, rigorous evaluation benchmarks and metrics to accurately measure lip synchronization in unconstrained videos. Extensive quantitative evaluations on our challenging benchmarks show that the lip-sync accuracy of the videos generated by our Wav2Lip model is almost as good as real synced videos. We provide a demo video clearly showing the substantial impact of our Wav2Lip model and evaluation benchmarks on our website: \url{cvit.iiit.ac.in/research/projects/cvit-projects/a-lip-sync-expert-is-all-you-need-for-speech-to-lip-generation-in-the-wild}. The code and models are released at this GitHub repository: \url{github.com/Rudrabha/Wav2Lip}. You can also try out the interactive demo at this link: \url{bhaasha.iiit.ac.in/lipsync}.
9 pages (including references), 3 figures, Accepted in ACM Multimedia, 2020
References in corpus (1)
Cited by in corpus (49)
- Deepfake Detection by Human Crowds, Machines, and Machine-informed Crowds
- Evaluation of an Audio-Video Multimodal Deepfake Dataset using Unimodal and Multimodal Detectors
- Emotional Speech-Driven Animation with Content-Emotion Disentanglement
- Audio-Driven Talking Face Video Generation with Dynamic Convolution Kernels
- Imitating Arbitrary Talking Style for Realistic Audio-DrivenTalking Face Synthesis
- Deep Person Generation: A Survey from the Perspective of Face, Pose and Cloth Synthesis
- DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion Autoencoder
- Towards Realistic Visual Dubbing with Heterogeneous Sources
- GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian Splatting
- Talking Head from Speech Audio using a Pre-trained Image Generator
- Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation
- FakeAVCeleb: A Novel Audio-Video Multimodal Deepfake Dataset
- AVTENet: A Human-Cognition-Inspired Audio-Visual Transformer-Based Ensemble Network for Video Deepfake Detection
- EmoFace: Audio-driven Emotional 3D Face Animation
- Lip-to-Speech Synthesis for Arbitrary Speakers in the Wild
- NeRF-AD: Neural Radiance Field with Attention-based Disentanglement for Talking Face Synthesis
- Audio-Driven Dubbing for User Generated Contents via Style-Aware Semi-Parametric Synthesis
- Human Motion Video Generation: A Survey
- LPIPS-AttnWav2Lip: Generic Audio-Driven lip synchronization for Talking Head Generation in the Wild
- Talking Head Generation with Audio and Speech Related Facial Action Units
- Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion
- Towards Accurate Lip-to-Speech Synthesis in-the-Wild
- DualVoice: Speech Interaction that Discriminates between Normal and Whispered Voice Input
- Memories are One-to-Many Mapping Alleviators in Talking Face Generation
- KoDF: A Large-scale Korean DeepFake Detection Dataset
- Write-a-speaker: Text-based Emotional and Rhythmic Talking-head Generation
- Intelligent Video Editing: Incorporating Modern Talking Face Generation Algorithms in a Video Editor
- Towards Attention-based Contrastive Learning for Audio Spoof Detection
- AV-Lip-Sync+: Leveraging AV-HuBERT to Exploit Multimodal Inconsistency for Deepfake Detection of Frontal Face Videos
- Steganography Beyond Space-Time with Chain of Multimodal AI
- Can deepfakes be created by novice users?
- Learning and Evaluating Human Preferences for Conversational Head Generation
- REFA: Real-time Egocentric Facial Animations for Virtual Reality
- Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head Synthesis
- NWT: Towards natural audio-to-video generation with representation learning
- Speech2Video: Cross-Modal Distillation for Speech to Video Generation
- Neural Dubber: Dubbing for Videos According to Scripts
- Generative AI for Video Translation: A Scalable Architecture for Multilingual Video Conferencing
- Impact of Benign Modifications on Discriminative Performance of Deepfake Detectors
- Evaluation of End-to-End Continuous Spanish Lipreading in Different Data Conditions
- Personalized One-Shot Lipreading for an ALS Patient
- Extreme-scale Talking-Face Video Upsampling with Audio-Visual Priors
- Visual Speech Enhancement Without A Real Visual Stream
- Advancing Talking Head Generation: A Comprehensive Survey of Multi-Modal Methodologies, Datasets, Evaluation Metrics, and Loss Functions
- A Comprehensive Multi-scale Approach for Speech and Dynamics Synchrony in Talking Head Generation
- AnyoneNet: Synchronized Speech and Talking Head Generation for Arbitrary Person
- Improving Generalization in Deepfake Detection with Face Foundation Models and Metric Learning
- MuteSwap: Visual-informed Silent Video Identity Conversion
- Multi-modal Speech Transformer Decoders: When Do Multiple Modalities Improve Accuracy?