A Lightweight Graph Transformer Network for Human Mesh Reconstruction from 2D Human Pose
arXiv:2111.12696
Abstract
Existing deep learning-based human mesh reconstruction approaches have a tendency to build larger networks in order to achieve higher accuracy. Computational complexity and model size are often neglected, despite being key characteristics for practical use of human mesh reconstruction models (e.g. virtual try-on systems). In this paper, we present GTRS, a lightweight pose-based method that can reconstruct human mesh from 2D human pose. We propose a pose analysis module that uses graph transformers to exploit structured and implicit joint correlations, and a mesh regression module that combines the extracted pose feature with the mesh template to reconstruct the final human mesh. We demonstrate the efficiency and generalization of GTRS by extensive evaluations on the Human3.6M and 3DPW datasets. In particular, GTRS achieves better accuracy than the SOTA pose-based method Pose2Mesh while only using 10.2% of the parameters (Params) and 2.5% of the FLOPs on the challenging in-the-wild 3DPW dataset. Code will be publicly available.
ACM Multimedia 2022
References in corpus (6)
- MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications
- Deformable DETR: Deformable Transformers for End-to-End Object Detection
- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
- GraFormer: Graph Convolution Transformer for 3D Pose Estimation
- 3D Human Pose Estimation with Spatial and Temporal Transformers
- Towards Fast and Accurate Multi-Person Pose Estimation on Mobile Devices