Towards Accurate Markerless Human Shape and Pose Estimation over Time
arXiv:1707.07548
Abstract
Existing marker-less motion capture methods often assume known backgrounds, static cameras, and sequence specific motion priors, which narrows its application scenarios. Here we propose a fully automatic method that given multi-view video, estimates 3D human motion and body shape. We take recent SMPLify \cite{bogo2016keep} as the base method, and extend it in several ways. First we fit the body to 2D features detected in multi-view images. Second, we use a CNN method to segment the person in each image and fit the 3D body model to the contours to further improves accuracy. Third we utilize a generic and robust DCT temporal prior to handle the left and right side swapping issue sometimes introduced by the 2D pose estimator. Validation on standard benchmarks shows our results are comparable to the state of the art and also provide a realistic 3D shape avatar. We also demonstrate accurate results on HumanEva and on challenging dance sequences from YouTube in monocular case.
10 pages, 6 figures, 5 tables, published in 3DV-2017
Cited by in corpus (21)
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human Digitization
- Weakly-Supervised Discovery of Geometry-Aware Representation for 3D Human Pose Estimation
- MonoPerfCap: Human Performance Capture from Monocular Video
- MoSculp: Interactive Visualization of Shape and Time
- Learning to Estimate 3D Human Pose and Shape from a Single Color Image
- DeepCap: Monocular Human Performance Capture Using Weak Supervision
- Delving Deep into Pixel Alignment Feature for Accurate Multi-view Human Mesh Recovery
- PyMAF: 3D Human Pose and Shape Regression with Pyramidal Mesh Alignment Feedback Loop
- Learning 3D Human Dynamics from Video
- LiveCap: Real-time Human Performance Capture from Monocular Video
- We are More than Our Joints: Predicting how 3D Bodies Move
- Bilevel Online Adaptation for Out-of-Domain Human Mesh Reconstruction
- Task-Generic Hierarchical Human Motion Prior using VAEs
- Can Action be Imitated? Learn to Reconstruct and Transfer Human Dynamics from Videos
- ChallenCap: Monocular 3D Capture of Challenging Human Performances using Multi-Modal References
- PedX: Benchmark Dataset for Metric 3D Pose Estimation of Pedestrians in Complex Urban Intersections
- Fashion is Taking Shape: Understanding Clothing Preference Based on Body Shape From Online Sources
- TexturePose: Supervising Human Mesh Estimation with Texture Consistency
- FAKIR: An algorithm for revealing the anatomy and pose of statues from raw point sets
- Camera Motion Agnostic 3D Human Pose Estimation
- Birds of a Feather: Capturing Avian Shape Models from Images