Robust 3D Hand Pose Estimation in Single Depth Images: from Single-View CNN to Multi-View CNNs
arXiv:1606.07253
Abstract
Articulated hand pose estimation plays an important role in human-computer interaction. Despite the recent progress, the accuracy of existing methods is still not satisfactory, partially due to the difficulty of embedded high-dimensional and non-linear regression problem. Different from the existing discriminative methods that regress for the hand pose with a single depth image, we propose to first project the query depth image onto three orthogonal planes and utilize these multi-view projections to regress for 2D heat-maps which estimate the joint positions on each plane. These multi-view heat-maps are then fused to produce final 3D hand pose estimation with learned pose priors. Experiments show that the proposed method largely outperforms state-of-the-art on a challenging dataset. Moreover, a cross-dataset experiment also demonstrates the good generalization ability of the proposed method.
9 pages, 9 figures, published at Computer Vision and Pattern Recognition (CVPR) 2016
References in corpus (1)
Cited by in corpus (17)
- Real-time Hand Tracking under Occlusion from an Egocentric RGB-D Sensor
- Hand3D: Hand Pose Estimation using 3D Neural Network
- V2V-PoseNet: Voxel-to-Voxel Prediction Network for Accurate 3D Hand and Human Pose Estimation from a Single Depth Map
- 3D Hand Shape and Pose Estimation from a Single RGB Image
- BigHand2.2M Benchmark: Hand Pose Dataset and State of the Art Analysis
- BiHand: Recovering Hand Mesh with Multi-stage Bisected Hourglass Networks
- Graph Edge Convolutional Neural Networks for Skeleton Based Action Recognition
- Model-based 3D Hand Reconstruction via Self-Supervised Learning
- Simultaneous Hand Pose and Skeleton Bone-Lengths Estimation from a Single Depth Image
- GANerated Hands for Real-time 3D Hand Tracking from Monocular RGB
- DeepHPS: End-to-end Estimation of 3D Hand Pose and Shape by Learning from Synthetic Depth
- Point-to-Pose Voting based Hand Pose Estimation using Residual Permutation Equivariant Layer
- Deep Kinematic Pose Regression
- Monocular 3D Reconstruction of Interacting Hands via Collision-Aware Factorized Refinements
- Enhanced Touchable Projector-depth System with Deep Hand Pose Estimation
- End-to-end 3D face reconstruction with deep neural networks
- Human Recognition Using Face in Computed Tomography