Head2Head++: Deep Facial Attributes Re-Targeting
arXiv:2006.10199 · doi:10.1109/TBIOM.2021.3049576
Abstract
Facial video re-targeting is a challenging problem aiming to modify the facial attributes of a target subject in a seamless manner by a driving monocular sequence. We leverage the 3D geometry of faces and Generative Adversarial Networks (GANs) to design a novel deep learning architecture for the task of facial and head reenactment. Our method is different to purely 3D model-based approaches, or recent image-based methods that use Deep Convolutional Neural Networks (DCNNs) to generate individual frames. We manage to capture the complex non-rigid facial motion from the driving monocular performances and synthesise temporally consistent videos, with the aid of a sequential Generator and an ad-hoc Dynamics Discriminator network. We conduct a comprehensive set of quantitative and qualitative tests and demonstrate experimentally that our proposed method can successfully transfer facial expressions, head pose and eye gaze from a source video to a target subject, in a photo-realistic and faithful fashion, better than other state-of-the-art methods. Most importantly, our system performs end-to-end reenactment in nearly real-time speed (18 fps).
Published in IEEE Transactions on Biometrics, Behavior, and Identity Science (Volume: 3, Issue: 1, Jan. 2021)
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium
- Deep Convolutional Inverse Graphics Network
- Towards Accurate Generative Models of Video: A New Metric & Challenges
- First Order Motion Model for Image Animation
- X2Face: A network for controlling face generation by using images, audio, and pose codes
Cited by in corpus (4)
- AvatarMAV: Fast 3D Head Avatar Reconstruction Using Motion-Aware Neural Voxels
- LatentAvatar: Learning Latent Expression Code for Expressive Neural Head Avatar
- Deep Person Generation: A Survey from the Perspective of Face, Pose and Cloth Synthesis
- Head2HeadFS: Video-based Head Reenactment with Few-shot Learning