Human Pose Regression by Combining Indirect Part Detection and Contextual Information
arXiv:1710.02322
Abstract
In this paper, we propose an end-to-end trainable regression approach for human pose estimation from still images. We use the proposed Soft-argmax function to convert feature maps directly to joint coordinates, resulting in a fully differentiable framework. Our method is able to learn heat maps representations indirectly, without additional steps of artificial ground truth generation. Consequently, contextual information can be included to the pose predictions in a seamless way. We evaluated our method on two very challenging datasets, the Leeds Sports Poses (LSP) and the MPII Human Pose datasets, reaching the best performance among all the existing regression methods and comparable results to the state-of-the-art detection based approaches.
References in corpus (4)
Cited by in corpus (6)
- Multi-task Deep Learning for Real-Time 3D Human Pose Estimation and Action Recognition
- Rethinking on Multi-Stage Networks for Human Pose Estimation
- Simple and Lightweight Human Pose Estimation
- DeepFuse: An IMU-Aware Network for Real-Time 3D Human Pose Estimation from Multi-View Image
- DeepSkeleton: Skeleton Map for 3D Human Pose Regression
- Anti-Confusing: Region-Aware Network for Human Pose Estimation