Geometric Pose Affordance: 3D Human Pose with Scene Constraints
arXiv:1905.07718
Abstract
Full 3D estimation of human pose from a single image remains a challenging task despite many recent advances. In this paper, we explore the hypothesis that strong prior information about scene geometry can be used to improve pose estimation accuracy. To tackle this question empirically, we have assembled a novel dataset, consisting of multi-view imagery of people interacting with a variety of rich 3D environments. We utilized a commercial motion capture system to collect gold-standard estimates of pose and construct accurate geometric 3D CAD models of the scene itself. To inject prior knowledge of scene constraints into existing frameworks for pose estimation from images, we introduce a novel, view-based representation of scene geometry, a , which employs multi-hit ray tracing to concisely encode multiple surface entry and exit points along each camera view ray direction. We propose two different mechanisms for integrating multi-layer depth information pose estimation: input as encoded ray features used in lifting 2D pose to full 3D, and secondly as a differentiable loss that encourages learned models to favor geometrically consistent pose estimates. We show experimentally that these techniques can improve the accuracy of 3D pose estimates, particularly in the presence of occlusion and complex scene geometry.
, in submission to CVIU
Cited by in corpus (15)
- TransFusion: Cross-view Fusion with Transformer for 3D Human Pose Estimation
- Adaptive Graphical Model Network for 2D Handpose Estimation
- Generating Realistic Training Images Based on Tonality-Alignment Generative Adversarial Networks for Hand Pose Estimation
- Placing Human Animations into 3D Scenes by Learning Interaction- and Geometry-Driven Keyframes
- Long-term Human Motion Prediction with Scene Context
- Hand-Object Contact Consistency Reasoning for Human Grasps Generation
- Predicting Camera Viewpoint Improves Cross-dataset Generalization for 3D Human Pose Estimation
- Synthesizing Long-Term 3D Human Motion and Interaction in 3D Scenes
- Rotation-invariant Mixed Graphical Model Network for 2D Hand Pose Estimation
- PACE: Data-Driven Virtual Agent Interaction in Dense and Cluttered Environments
- Knowledge Integration Networks for Action Recognition
- PC-HMR: Pose Calibration for 3D Human Mesh Recovery from 2D Images/Videos
- SGE net: Video object detection with squeezed GRU and information entropy map
- Learning Local Recurrent Models for Human Mesh Recovery
- DGGAN: Depth-image Guided Generative Adversarial Networks for Disentangling RGB and Depth Images in 3D Hand Pose Estimation