KinePose: A temporally optimized inverse kinematics technique for 6DOF human pose estimation with biomechanical constraints
arXiv:2207.12841 · doi:10.56541/QTUV2945
Abstract
Computer vision/deep learning-based 3D human pose estimation methods aim to localize human joints from images and videos. Pose representation is normally limited to 3D joint positional/translational degrees of freedom (3DOFs), however, a further three rotational DOFs (6DOFs) are required for many potential biomechanical applications. Positional DOFs are insufficient to analytically solve for joint rotational DOFs in a 3D human skeletal model. Therefore, we propose a temporal inverse kinematics (IK) optimization technique to infer joint orientations throughout a biomechanically informed, and subject-specific kinematic chain. For this, we prescribe link directions from a position-based 3D pose estimate. Sequential least squares quadratic programming is used to solve a minimization problem that involves both frame-based pose terms, and a temporal term. The solution space is constrained using joint DOFs, and ranges of motion (ROMs). We generate 3D pose motion sequences to assess the IK approach both for general accuracy, and accuracy in boundary cases. Our temporal algorithm achieves 6DOF pose estimates with low Mean Per Joint Angular Separation (MPJAS) errors (3.7°/joint overall, & 1.6°/joint for lower limbs). With frame-by-frame IK we obtain low errors in the case of bent elbows and knees, however, motion sequences with phases of extended/straight limbs results in ambiguity in twist angle. With temporal IK, we reduce ambiguity for these poses, resulting in lower average errors.
Project page: https://kevgildea.github.io/KinePose/
References in corpus (27)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Bootstrap your own latent: A new approach to self-supervised Learning
- Learning-Based View Synthesis for Light Field Cameras
- Learning Representations by Maximizing Mutual Information Across Views
- Monocular Human Pose Estimation: A Survey of Deep Learning-based Methods
- Do ImageNet Classifiers Generalize to ImageNet?
- Sparsity in Deep Learning: Pruning and growth for efficient inference and training in neural networks
- Lung and Colon Cancer Histopathological Image Dataset (LC25000)
- A new face database simultaneously acquired in visible, near infrared and thermal spectrum
- Local Light Field Fusion: Practical View Synthesis with Prescriptive Sampling Guidelines
- Convolution Neural Networks for diagnosing colon and lung cancer histopathological images
- Video Object Segmentation with Re-identification
- Point-wise Map Recovery and Refinement from Functional Correspondence
- Self-Correction for Human Parsing
- Model Patching: Closing the Subgroup Performance Gap with Data Augmentation
- RVOS: End-to-End Recurrent Network for Video Object Segmentation
- Fast and Accurate Optical Flow based Depth Map Estimation from Light Fields
- CamSwarm: Instantaneous Smartphone Camera Arrays for Collaborative Photography
- Pruning Convolutional Neural Networks with Self-Supervision
- ChipNet: Budget-Aware Pruning with Heaviside Continuous Approximations
- Temporally Consistent Video Colorization with Deep Feature Propagation and Self-regularization Learning
- PoseFace: Pose-Invariant Features and Pose-Adaptive Loss for Face Recognition
- Towards End-to-End Neural Face Authentication in the Wild -- Quantifying and Compensating for Directional Lighting Effects
- Edge-aware Bidirectional Diffusion for Dense Depth Estimation from Light Fields
- A Study of Efficient Light Field Subsampling and Reconstruction Strategies
- Shape Analysis via Functional Map Construction and Bases Pursuit