SparsePoser: Real-time Full-body Motion Reconstruction from Sparse Data
arXiv:2311.02191 · doi:10.1145/3625264
Abstract
Accurate and reliable human motion reconstruction is crucial for creating natural interactions of full-body avatars in Virtual Reality (VR) and entertainment applications. As the Metaverse and social applications gain popularity, users are seeking cost-effective solutions to create full-body animations that are comparable in quality to those produced by commercial motion capture systems. In order to provide affordable solutions, though, it is important to minimize the number of sensors attached to the subject's body. Unfortunately, reconstructing the full-body pose from sparse data is a heavily under-determined problem. Some studies that use IMU sensors face challenges in reconstructing the pose due to positional drift and ambiguity of the poses. In recent years, some mainstream VR systems have released 6-degree-of-freedom (6-DoF) tracking devices providing positional and rotational information. Nevertheless, most solutions for reconstructing full-body poses rely on traditional inverse kinematics (IK) solutions, which often produce non-continuous and unnatural poses. In this article, we introduce SparsePoser, a novel deep learning-based solution for reconstructing a full-body pose from a reduced set of six tracking devices. Our system incorporates a convolutional-based autoencoder that synthesizes high-quality continuous human poses by learning the human motion manifold from motion capture data. Then, we employ a learned IK component, made of multiple lightweight feed-forward neural networks, to adjust the hands and feet toward the corresponding trackers. We extensively evaluate our method on publicly available motion capture datasets and with real-time live demos. We show that our method outperforms state-of-the-art techniques using IMU sensors or 6-DoF tracking devices, and can be used for users with different body dimensions and proportions.
Published in ACM TOG https://dl.acm.org/doi/10.1145/3625264 and presented in SIGGRAPH ASIA 2023
References in corpus (9)
- Skeleton-Aware Networks for Deep Motion Retargeting
- QuestSim: Human Motion Tracking from Sparse Sensors with Simulated Avatars
- Transformer Inertial Poser: Real-time Human Motion Reconstruction from Sparse IMUs with Simultaneous Terrain Generation
- HUMAN4D: A Human-Centric Multimodal Dataset for Motions and Immersive Media
- Combining Motion Matching and Orientation Prediction to Animate Avatars for Consumer-Grade VR Devices
- Animation Fidelity in Self-Avatars: Impact on User Performance and Sense of Agency
- Pose Representations for Deep Skeletal Animation
- Learning-based pose edition for efficient and interactive design
- Neural Inverse Kinematics
Cited by in corpus (5)
- MobilePoser: Real-Time Full-Body Pose Estimation and 3D Human Translation from IMUs in Mobile Consumer Devices
- Deep learning for 3D human pose estimation and mesh recovery: A survey
- ELMO: Enhanced Real-time LiDAR Motion Capture through Upsampling
- DragPoser: Motion Reconstruction from Variable Sparse Tracking Signals via Latent Space Optimization
- Physical Self-Supervised Learning: IMU Sensing without Manual Labels