Learning to Fuse 2D and 3D Image Cues for Monocular Body Pose Estimation
arXiv:1611.05708
Abstract
Most recent approaches to monocular 3D human pose estimation rely on Deep Learning. They typically involve regressing from an image to either 3D joint coordinates directly or 2D joint locations from which 3D coordinates are inferred. Both approaches have their strengths and weaknesses and we therefore propose a novel architecture designed to deliver the best of both worlds by performing both simultaneously and fusing the information along the way. At the heart of our framework is a trainable fusion scheme that learns how to fuse the information optimally instead of being hand-designed. This yields significant improvements upon the state-of-the-art on standard 3D human pose estimation benchmarks.
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Stacked Hourglass Networks for Human Pose Estimation
- MoCap-guided Data Augmentation for 3D Pose Estimation in the Wild
- Synthesizing Training Images for Boosting Human 3D Pose Estimation
- Structured Prediction of 3D Human Pose with Deep Neural Networks
- Coarse-to-Fine Volumetric Prediction for Single-Image 3D Human Pose
Cited by in corpus (10)
- VNect: Real-time 3D Human Pose Estimation with a Single RGB Camera
- Lifting from the Deep: Convolutional 3D Pose Estimation from a Single Image
- Fast and Robust Multi-Person 3D Pose Estimation from Multiple Views
- Learning to Estimate 3D Human Pose and Shape from a Single Color Image
- Monocular 3D Human Pose Estimation In The Wild Using Improved CNN Supervision
- DeepSkeleton: Skeleton Map for 3D Human Pose Regression
- Ordinal Depth Supervision for 3D Human Pose Estimation
- Rethinking Pose in 3D: Multi-stage Refinement and Recovery for Markerless Motion Capture
- Deep Autoencoder for Combined Human Pose Estimation and body Model Upscaling
- Coherent Reconstruction of Multiple Humans from a Single Image