3D Human Texture Estimation from a Single Image with Transformers
arXiv:2109.02563
Abstract
We propose a Transformer-based framework for 3D human texture estimation from a single image. The proposed Transformer is able to effectively exploit the global information of the input image, overcoming the limitations of existing methods that are solely based on convolutional neural networks. In addition, we also propose a mask-fusion strategy to combine the advantages of the RGB-based and texture-flow-based models. We further introduce a part-style loss to help reconstruct high-fidelity colors without introducing unpleasant artifacts. Extensive experiments demonstrate the effectiveness of the proposed method against state-of-the-art 3D human texture estimation approaches both quantitatively and qualitatively.
ICCV 2021 Oral, Project: https://www.mmlab-ntu.com/project/texformer, Code: https://github.com/xuxy09/Texformer
References in corpus (10)
- Beyond Part Models: Person Retrieval with Refined Part Pooling (and a Strong Convolutional Baseline)
- Torchreid: A Library for Deep Learning Person Re-Identification in Pytorch
- Learning Spatial and Spatio-Temporal Pixel Aggregations for Image and Video Denoising
- Quadratic video interpolation
- EANet: Enhancing Alignment for Cross-Domain Person Re-identification
- Learning to Reconstruct People in Clothing from a Single RGB Camera
- ARCH: Animatable Reconstruction of Clothed Humans
- TexMesh: Reconstructing Detailed Human Texture and Geometry from RGB-D Video
- Re-Identification Supervised Texture Generation
- Reconstructing NBA Players