Lightweight Multi-person Total Motion Capture Using Sparse Multi-view Cameras
arXiv:2108.10378
Abstract
Multi-person total motion capture is extremely challenging when it comes to handle severe occlusions, different reconstruction granularities from body to face and hands, drastically changing observation scales and fast body movements. To overcome these challenges above, we contribute a lightweight total motion capture system for multi-person interactive scenarios using only sparse multi-view cameras. By contributing a novel hand and face bootstrapping algorithm, our method is capable of efficient localization and accurate association of the hands and faces even on severe occluded occasions. We leverage both pose regression and keypoints detection methods and further propose a unified two-stage parametric fitting method for achieving pixel-aligned accuracy. Moreover, for extremely self-occluded poses and close interactions, a novel feedback mechanism is proposed to propagate the pixel-aligned reconstructions into the next frame for more accurate association. Overall, we propose the first light-weight total capture system and achieves fast, robust and accurate multi-person total motion capture performance. The results and experiments show that our method achieves more accurate results than existing methods under sparse-view setups.
References in corpus (7)
- FrankMocap: Fast Monocular 3D Hand and Body Motion Capture by Regression and Integration
- Fast and Robust Multi-Person 3D Pose Estimation from Multiple Views
- Panoptic Studio: A Massively Multiview System for Social Interaction Capture
- 4D Association Graph for Realtime Multi-person Motion Capture Using Multiple Video Cameras
- Monocular Real-time Full Body Capture with Inter-part Correlations
- POSEFusion: Pose-guided Selective Fusion for Single-view Human Volumetric Capture
- Neural Free-Viewpoint Performance Rendering under Complex Human-object Interactions