Deep Spatial Transformation for Pose-Guided Person Image Generation and Animation
arXiv:2008.12606 · doi:10.1109/TIP.2020.3018224
Abstract
Pose-guided person image generation and animation aim to transform a source person image to target poses. These tasks require spatial manipulation of source data. However, Convolutional Neural Networks are limited by the lack of ability to spatially transform the inputs. In this paper, we propose a differentiable global-flow local-attention framework to reassemble the inputs at the feature level. This framework first estimates global flow fields between sources and targets. Then, corresponding local source feature patches are sampled with content-aware local attention coefficients. We show that our framework can spatially transform the inputs in an efficient manner. Meanwhile, we further model the temporal consistency for the person image animation task to generate coherent videos. The experiment results of both image generation and animation tasks demonstrate the superiority of our model. Besides, additional results of novel view synthesis and face image animation show that our model is applicable to other tasks requiring spatial transformation. The source code of our project is available at https://github.com/RenYurui/Global-Flow-Local-Attention.
arXiv admin note: text overlap with arXiv:2003.00696
References in corpus (3)
Cited by in corpus (11)
- Deep Person Generation: A Survey from the Perspective of Face, Pose and Cloth Synthesis
- F3A-GAN: Facial Flow for Face Animation with Generative Adversarial Networks
- PISE: Person Image Synthesis and Editing with Decoupled GAN
- PIRenderer: Controllable Portrait Image Generation via Semantic Neural Rendering
- Human Image Generation: A Comprehensive Survey
- An Identity-Preserved Framework for Human Motion Transfer
- Shape-Guided Clothing Warping for Virtual Try-On
- Flow Guided Transformable Bottleneck Networks for Motion Retargeting
- HumanRAM: Feed-forward Human Reconstruction and Animation Model using Transformers
- A 3D Mesh-based Lifting-and-Projection Network for Human Pose Transfer
- GLocal: Global Graph Reasoning and Local Structure Transfer for Person Image Generation