Capturing Hands in Action using Discriminative Salient Points and Physics Simulation
arXiv:1506.02178 · doi:10.1007/s11263-016-0895-4
Abstract
Hand motion capture is a popular research field, recently gaining more attention due to the ubiquity of RGB-D sensors. However, even most recent approaches focus on the case of a single isolated hand. In this work, we focus on hands that interact with other hands or objects and present a framework that successfully captures motion in such interaction scenarios for both rigid and articulated objects. Our framework combines a generative model with discriminatively trained salient points to achieve a low tracking error and with collision detection and physics simulation to achieve physically plausible estimates even in case of occlusions and missing visual data. Since all components are unified in a single objective function which is almost everywhere differentiable, it can be optimized with standard optimization techniques. Our approach works for monocular RGB-D sequences as well as setups with multiple synchronized RGB cameras. For a qualitative and quantitative evaluation, we captured 29 sequences with a large variety of interactions and up to 150 degrees of freedom.
Accepted for publication by the International Journal of Computer Vision (IJCV) on 16.02.2016 (submitted on 17.10.14). A combination into a single framework of an ECCV'12 multicamera-RGB and a monocular-RGBD GCPR'14 hand tracking paper with several extensions, additional experiments and details
References in corpus (4)
Cited by in corpus (43)
- Embodied Hands: Modeling and Capturing Hands and Bodies Together
- Real-time Hand Tracking under Occlusion from an Egocentric RGB-D Sensor
- GRAB: A Dataset of Whole-Body Human Grasping of Objects
- Real-time Pose and Shape Reconstruction of Two Interacting Hands With a Single Depth Camera
- SignBERT+: Hand-model-aware Self-supervised Pre-training for Sign Language Understanding
- Joint Hand-object 3D Reconstruction from a Single Image with Cross-branch Feature Fusion
- FrankMocap: Fast Monocular 3D Hand and Body Motion Capture by Regression and Integration
- RGB2Hands: Real-Time Tracking of 3D Hand Interactions from Monocular RGB Video
- Hand Pose Estimation: A Survey
- Learning Multi-Human Optical Flow
- Real-time Joint Tracking of a Hand Manipulating an Object from RGB-D Input
- Survey on Hand Gesture Recognition from Visual Input
- Monocular Total Capture: Posing Face, Body, and Hands in the Wild
- First-Person Hand Action Benchmark with RGB-D Videos and 3D Hand Pose Annotations
- Physical Interaction: Reconstructing Hand-object Interactions with Physics
- POV-Surgery: A Dataset for Egocentric Hand and Tool Pose Estimation During Surgical Activities
- Disentangling Latent Hands for Image Synthesis and Pose Estimation
- Grasping Field: Learning Implicit Representations for Human Grasps
- Occlusion-aware Hand Pose Estimation Using Hierarchical Mixture Density Network
- Physics-Based Dexterous Manipulations with Estimated Hand Poses and Residual Reinforcement Learning
- Monocular Real-time Hand Shape and Motion Capture using Multi-modal Data
- Pyramid Deep Fusion Network for Two-Hand Reconstruction from RGB-D Images
- GANerated Hands for Real-time 3D Hand Tracking from Monocular RGB
- Hand-Object Contact Consistency Reasoning for Human Grasps Generation
- Perceiving 3D Human-Object Spatial Arrangements from a Single Image in the Wild
- Task-Oriented Hand Motion Retargeting for Dexterous Manipulation Imitation
- EasyHOI: Unleashing the Power of Large Models for Reconstructing Hand-Object Interactions in the Wild
- JGR-P2O: Joint Graph Reasoning based Pixel-to-Offset Prediction Network for 3D Hand Pose Estimation from a Single Depth Image
- Back to RGB: 3D tracking of hands and hand-object interactions based on short-baseline stereo
- CPF: Learning a Contact Potential Field to Model the Hand-Object Interaction
- ClickDiff: Click to Induce Semantic Contact Map for Controllable Grasp Generation with Diffusion Models
- Two-hand Global 3D Pose Estimation Using Monocular RGB
- Leveraging Photometric Consistency over Time for Sparsely Supervised Hand-Object Reconstruction
- Unsupervised Domain Adaptation with Temporal-Consistent Self-Training for 3D Hand-Object Joint Reconstruction
- Monocular 3D Reconstruction of Interacting Hands via Collision-Aware Factorized Refinements
- A Modular Approach to the Embodiment of Hand Motions from Human Demonstrations
- SignBERT: Pre-Training of Hand-Model-Aware Representation for Sign Language Recognition
- Physically Plausible Pose Refinement using Fully Differentiable Forces
- Holistic 3D Human and Scene Mesh Estimation from Single View Images
- Hand-Shadow Poser
- Parallel mesh reconstruction streams for pose estimation of interacting hands
- Learning to Train with Synthetic Humans
- Towards Markerless Grasp Capture