Deep Person Generation: A Survey from the Perspective of Face, Pose and Cloth Synthesis
arXiv:2109.02081 · doi:10.1145/3575656
Abstract
Deep person generation has attracted extensive research attention due to its wide applications in virtual agents, video conferencing, online shopping and art/movie production. With the advancement of deep learning, visual appearances (face, pose, cloth) of a person image can be easily generated or manipulated on demand. In this survey, we first summarize the scope of person generation, and then systematically review recent progress and technical trends in deep person generation, covering three major tasks: talking-head generation (face), pose-guided person generation (pose) and garment-oriented person generation (cloth). More than two hundred papers are covered for a thorough overview, and the milestone works are highlighted to witness the major technical breakthrough. Based on these fundamental tasks, a number of applications are investigated, e.g., virtual fitting, digital human, generative data augmentation. We hope this survey could shed some light on the future prospects of deep person generation, and provide a helpful foundation for full applications towards digital human.
References in corpus (18)
- Distilling the Knowledge in a Neural Network
- Conditional Generative Adversarial Nets
- A Lip Sync Expert Is All You Need for Speech to Lip Generation In The Wild
- Fast Parallel Hypertree Decompositions in Logarithmic Recursion Depth
- ObamaNet: Photo-realistic lip-sync from text
- Audio-driven Talking Face Video Generation with Learning-based Personalized Head Pose
- Deep Spatial Transformation for Pose-Guided Person Image Generation and Animation
- PoNA: Pose-guided Non-local Attention for Human Pose Transfer
- End-to-End Learning of Geometric Deformations of Feature Maps for Virtual Try-On
- Shape Controllable Virtual Try-on for Underwear Models
- Unsupervised Pose Flow Learning for Pose Guided Synthesis
- FaR-GAN for One-Shot Face Reenactment
- DwNet: Dense warp-based network for pose-guided human video generation
- Speech Driven Talking Face Generation from a Single Image and an Emotion Condition
- Pose-Guided High-Resolution Appearance Transfer via Progressive Training
- Robust Pose Transfer with Dynamic Details using Neural Video Rendering
- Everything's Talkin': Pareidolia Face Reenactment
- Person image generation with semantic attention network for person re-identification