Embodied Hands: Modeling and Capturing Hands and Bodies Together
arXiv:2201.02610 · doi:10.1145/3130800.3130883
Abstract
Humans move their hands and bodies together to communicate and solve tasks. Capturing and replicating such coordinated activity is critical for virtual characters that behave realistically. Surprisingly, most methods treat the 3D modeling and tracking of bodies and hands separately. Here we formulate a model of hands and bodies interacting together and fit it to full-body 4D sequences. When scanning or capturing the full body in 3D, hands are small and often partially occluded, making their shape and pose hard to recover. To cope with low-resolution, occlusion, and noise, we develop a new model called MANO (hand Model with Articulated and Non-rigid defOrmations). MANO is learned from around 1000 high-resolution 3D scans of hands of 31 subjects in a wide variety of hand poses. The model is realistic, low-dimensional, captures non-rigid shape changes with pose, is compatible with standard graphics packages, and can fit any human hand. MANO provides a compact mapping from hand poses to pose blend shape corrections and a linear manifold of pose synergies. We attach MANO to a standard parameterized 3D body shape model (SMPL), resulting in a fully articulated body and hand model (SMPL+H). We illustrate SMPL+H by fitting complex, natural, activities of subjects captured with a 4D scanner. The fitting is fully automatic and results in full body models that move naturally with detailed hand motions and a realism not seen before in full body performance capture. The models and data are freely available for research purposes in our website (http://mano.is.tue.mpg.de).
SIGGRAPH ASIA 2017
Cited by in corpus (57)
- GRAB: A Dataset of Whole-Body Human Grasping of Objects
- Real-time Pose and Shape Reconstruction of Two Interacting Hands With a Single Depth Camera
- Differentiable Rendering: A Survey
- Modeling Clothing as a Separate Layer for an Animatable Human Avatar
- FrankMocap: Fast Monocular 3D Hand and Body Motion Capture by Regression and Integration
- 3D Hand Shape and Pose Estimation from a Single RGB Image
- History Repeats Itself: Human Motion Prediction via Motion Attention
- HandAugment: A Simple Data Augmentation Method for Depth-Based 3D Hand Pose Estimation
- GraFormer: Graph Convolution Transformer for 3D Pose Estimation
- Learning Object Manipulation Skills via Approximate State Estimation from Real Videos
- The Whole Is Greater Than the Sum of Its Nonrigid Parts
- Recent Advances in 3D Object and Hand Pose Estimation
- DexPilot: Vision Based Teleoperation of Dexterous Robotic Hand-Arm System
- Physics-Based Dexterous Manipulations with Estimated Hand Poses and Residual Reinforcement Learning
- SimulCap : Single-View Human Performance Capture with Cloth Simulation
- Adversarial Motion Modelling helps Semi-supervised Hand Pose Estimation
- Model-based 3D Hand Reconstruction via Self-Supervised Learning
- Understanding Human Hands in Contact at Internet Scale
- We are More than Our Joints: Predicting how 3D Bodies Move
- Monocular Real-time Full Body Capture with Inter-part Correlations
- Hand-Object Contact Consistency Reasoning for Human Grasps Generation
- SCANimate: Weakly Supervised Learning of Skinned Clothed Avatar Networks
- Towards Accurate Alignment in Real-time 3D Hand-Mesh Reconstruction
- Two-hand Global 3D Pose Estimation Using Monocular RGB
- Whole-Body Human Pose Estimation in the Wild
- Monocular Real-Time Volumetric Performance Capture
- SeqHAND:RGB-Sequence-Based 3D Hand Pose and Shape Estimation
- Temporal-Aware Self-Supervised Learning for 3D Hand Pose and Mesh Estimation in Videos
- Unsupervised Shape and Pose Disentanglement for 3D Meshes
- Creating a Large-scale Synthetic Dataset for Human Activity Recognition
- Leveraging Photometric Consistency over Time for Sparsely Supervised Hand-Object Reconstruction
- Lightweight Multi-person Total Motion Capture Using Sparse Multi-view Cameras
- S3: Neural Shape, Skeleton, and Skinning Fields for 3D Human Modeling
- Monocular 3D Reconstruction of Interacting Hands via Collision-Aware Factorized Refinements
- Unsupervised Domain Adaptation with Temporal-Consistent Self-Training for 3D Hand-Object Joint Reconstruction
- Identity-Disentangled Neural Deformation Model for Dynamic Meshes
- Motion Capture from Internet Videos
- Reconstructing NBA Players
- MVHM: A Large-Scale Multi-View Hand Mesh Benchmark for Accurate 3D Hand Pose Estimation
- SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language Recognition
- HO-3D_v3: Improving the Accuracy of Hand-Object Annotations of the HO-3D Dataset
- Semi-Supervised 3D Hand-Object Poses Estimation with Interactions in Time
- DC-GNet: Deep Mesh Relation Capturing Graph Convolution Network for 3D Human Shape Reconstruction
- Parallel mesh reconstruction streams for pose estimation of interacting hands
- Hierarchical Neural Implicit Pose Network for Animation and Motion Retargeting
- LBS Autoencoder: Self-supervised Fitting of Articulated Meshes to Point Clouds
- Synthetic Training for Monocular Human Mesh Recovery
- A novel shape matching descriptor for real-time hand gesture recognition
- Learning Transferable Kinematic Dictionary for 3D Human Pose and Shape Reconstruction
- Towards Markerless Grasp Capture
- Single RGB-D Camera Teleoperation for General Robotic Manipulation
- Real-time RGBD-based Extended Body Pose Estimation
- Deep 3D Mesh Watermarking with Self-Adaptive Robustness
- Egocentric View Hand Action Recognition by Leveraging Hand Surface and Hand Grasp Type
- Birds of a Feather: Capturing Avian Shape Models from Images
- From Real to Synthetic and Back: Synthesizing Training Data for Multi-Person Scene Understanding
- TexturePose: Supervising Human Mesh Estimation with Texture Consistency