Maximum-Margin Structured Learning with Deep Networks for 3D Human Pose Estimation
arXiv:1508.06708
Abstract
This paper focuses on structured-output learning using deep neural networks for 3D human pose estimation from monocular images. Our network takes an image and 3D pose as inputs and outputs a score value, which is high when the image-pose pair matches and low otherwise. The network structure consists of a convolutional neural network for image feature extraction, followed by two sub-networks for transforming the image features and pose into a joint embedding. The score function is then the dot-product between the image and pose embeddings. The image-pose embedding and score function are jointly trained using a maximum-margin cost function. Our proposed framework can be interpreted as a special form of structured support vector machines where the joint feature space is discriminatively learned using deep neural networks. We test our framework on the Human3.6m dataset and obtain state-of-the-art results compared to other recent methods. Finally, we present visualizations of the image-pose embedding space, demonstrating the network has learned a high-level embedding of body-orientation and pose-configuration.
References in corpus (6)
- Theano: new features and speed improvements
- CNN Features off-the-shelf: an Astounding Baseline for Recognition
- Better Mixing via Deep Representations
- Learning Human Pose Estimation Features with Convolutional Networks
- Deep Structured Output Learning for Unconstrained Text Recognition
- Deep Structured learning for mass segmentation from Mammograms
Cited by in corpus (10)
- VNect: Real-time 3D Human Pose Estimation with a Single RGB Camera
- VSE++: Improving Visual-Semantic Embeddings with Hard Negatives
- MoCap-guided Data Augmentation for 3D Pose Estimation in the Wild
- MonoPerfCap: Human Performance Capture from Monocular Video
- 3D Human Pose Estimation from a Single Image via Distance Matrix Regression
- Deep Kinematic Pose Regression
- Human Motion Capture Using a Drone
- Single-Shot Multi-Person 3D Pose Estimation From Monocular RGB
- Deep Autoencoder for Combined Human Pose Estimation and body Model Upscaling
- 3D Human Pose Estimation Using Convolutional Neural Networks with 2D Pose Information