ObamaNet: Photo-realistic lip-sync from text
arXiv:1801.01442
Abstract
We present ObamaNet, the first architecture that generates both audio and synchronized photo-realistic lip-sync videos from any new text. Contrary to other published lip-sync approaches, ours is only composed of fully trainable neural modules and does not rely on any traditional computer graphics methods. More precisely, we use three main modules: a text-to-speech network based on Char2Wav, a time-delayed LSTM to generate mouth-keypoints synced to the audio, and a network based on Pix2Pix to generate the video frames conditioned on the keypoints.
Cited by in corpus (11)
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In The Wild
- Deep Audio-Visual Learning: A Survey
- Lip-to-Speech Synthesis for Arbitrary Speakers in the Wild
- A Survey and Taxonomy of Adversarial Neural Networks for Text-to-Image Synthesis
- LumièreNet: Lecture Video Synthesis from Audio
- Speaker Independent and Multilingual/Mixlingual Speech-Driven Talking Head Generation Using Phonetic Posteriorgrams
- Animating Face using Disentangled Audio Representations
- Iterative Text-based Editing of Talking-heads Using Neural Retargeting
- Visual Speech Enhancement Without A Real Visual Stream
- Lets Play Music: Audio-driven Performance Video Generation
- AnyoneNet: Synchronized Speech and Talking Head Generation for Arbitrary Person