VNect: Real-time 3D Human Pose Estimation with a Single RGB Camera
arXiv:1705.01583 · doi:10.1145/3072959.3073596
Abstract
We present the first real-time method to capture the full global 3D skeletal pose of a human in a stable, temporally consistent manner using a single RGB camera. Our method combines a new convolutional neural network (CNN) based pose regressor with kinematic skeleton fitting. Our novel fully-convolutional pose formulation regresses 2D and 3D joint positions jointly in real time and does not require tightly cropped input frames. A real-time kinematic skeleton fitting method uses the CNN output to yield temporally stable 3D global pose reconstructions on the basis of a coherent kinematic skeleton. This makes our approach the first monocular RGB method usable in real-time applications such as 3D character control---thus far, the only monocular methods for such applications employed specialized RGB-D cameras. Our method's accuracy is quantitatively on par with the best offline 3D monocular RGB pose estimation methods. Our results are qualitatively comparable to, and sometimes better than, results from monocular RGB-D approaches, such as the Kinect. However, we show that our approach is more broadly applicable than RGB-D solutions, i.e. it works for outdoor scenes, community videos, and low quality commodity RGB cameras.
Accepted to SIGGRAPH 2017
References in corpus (5)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- ADADELTA: An Adaptive Learning Rate Method
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Joint Training of a Convolutional Network and a Graphical Model for Human Pose Estimation
- MoCap-guided Data Augmentation for 3D Pose Estimation in the Wild
Cited by in corpus (166)
- All One Needs to Know about Metaverse: A Complete Survey on Technological Singularity, Virtual Ecosystem, and Research Agenda
- OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields
- Monocular Human Pose Estimation: A Survey of Deep Learning-based Methods
- Exploiting temporal information for 3D pose estimation
- XNect: Real-time Multi-Person 3D Motion Capture with a Single RGB Camera
- Unpaired Motion Style Transfer from Video to Animation
- Cascaded deep monocular 3D human pose estimation with evolutionary training data
- Multi-task Deep Learning for Real-Time 3D Human Pose Estimation and Action Recognition
- Subdivision-Based Mesh Convolution Networks
- Discovery of Latent 3D Keypoints via End-to-end Geometric Reasoning
- Rhythmic Gesticulator: Rhythm-Aware Co-Speech Gesture Synthesis with Hierarchical Neural Embeddings
- MotioNet: 3D Human Motion Reconstruction from Monocular Video with Skeleton Consistency
- MeTRAbs: Metric-Scale Truncation-Robust Heatmaps for Absolute 3D Human Pose Estimation
- IMUPoser: Full-Body Pose Estimation using IMUs in Phones, Watches, and Earbuds
- Vision-Based Human Pose Estimation via Deep Learning: A Survey
- Learning Character-Agnostic Motion for Motion Retargeting in 2D
- Deep Learning-Based Human Pose Estimation: A Survey
- Enabling Pedestrian Safety using Computer Vision Techniques: A Case Study of the 2018 Uber Inc. Self-driving Car Crash
- SelfPose: 3D Egocentric Pose Estimation from a Headset Mounted Camera
- BodyNet: Volumetric Inference of 3D Human Body Shapes
- Improving Robustness and Accuracy via Relative Information Encoding in 3D Human Pose Estimation
- FrankMocap: Fast Monocular 3D Hand and Body Motion Capture by Regression and Integration
- Deformation Capture via Soft and Stretchable Sensor Arrays
- Deep Part Induction from Articulated Object Pairs
- How Robust is 3D Human Pose Estimation to Occlusion?
- Neural Human Video Rendering by Learning Dynamic Textures and Rendering-to-Video Translation
- iMapper: Interaction-guided Joint Scene and Human Motion Mapping from Monocular Videos
- HUMAN4D: A Human-Centric Multimodal Dataset for Motions and Immersive Media
- MonoPerfCap: Human Performance Capture from Monocular Video
- Geometric Pose Affordance: 3D Human Pose with Scene Constraints
- Learning Dynamical Human-Joint Affinity for 3D Pose Estimation in Videos
- Single-shot 3D multi-person pose estimation in complex images
- Security Properties of Gait for Mobile Device Pairing
- Synthetic Occlusion Augmentation with Volumetric Heatmaps for the 2018 ECCV PoseTrack Challenge on 3D Human Pose Estimation
- Real-Time Human Pose Estimation on a Smart Walker using Convolutional Neural Networks
- TransFusion: Cross-view Fusion with Transformer for 3D Human Pose Estimation
- 3D Human Pose Estimation in the Wild by Adversarial Learning
- DeepMoCap: Deep Optical Motion Capture Using Multiple Depth Sensors and Retro-Reflectors
- Dual networks based 3D Multi-Person Pose Estimation from Monocular Video
- MoSculp: Interactive Visualization of Shape and Time
- Make Skeleton-based Action Recognition Model Smaller, Faster and Better
- Text as Neural Operator: Image Manipulation by Text Instruction
- Synergetic Reconstruction from 2D Pose and 3D Motion for Wide-Space Multi-Person Video Motion Capture in the Wild
- Monocular Total Capture: Posing Face, Body, and Hands in the Wild
- Video Based Reconstruction of 3D People Models
- Towards Single Camera Human 3D-Kinematics
- PhysCap: Physically Plausible Monocular 3D Motion Capture in Real Time
- Trajectory Space Factorization for Deep Video-Based 3D Human Pose Estimation
- DeepCap: Monocular Human Performance Capture Using Weak Supervision
- DeepHuman: 3D Human Reconstruction from a Single Image
- Mo2Cap2: Real-time Mobile 3D Motion Capture with a Cap-mounted Fisheye Camera
- Metric-Scale Truncation-Robust Heatmaps for 3D Human Pose Estimation
- SMPLR: Deep SMPL reverse for 3D human pose and shape recovery
- Inference Stage Optimization for Cross-scenario 3D Human Pose Estimation
- A-NeRF: Articulated Neural Radiance Fields for Learning Human Shape, Appearance, and Pose
- avaTTAR: Table Tennis Stroke Training with On-body and Detached Visualization in Augmented Reality
- MHFormer: Multi-Hypothesis Transformer for 3D Human Pose Estimation
- Skepxels: Spatio-temporal Image Representation of Human Skeleton Joints for Action Recognition
- SRNet: Improving Generalization in 3D Human Pose Estimation with a Split-and-Recombine Approach
- Holistic++ Scene Understanding: Single-view 3D Holistic Scene Parsing and Human Pose Estimation with Human-Object Interaction and Physical Commonsense
- Peeking into occluded joints: A novel framework for crowd pose estimation
- Drivable Avatar Clothing: Faithful Full-Body Telepresence with Dynamic Clothing Driven by Sparse RGB-D Input
- DenseRaC: Joint 3D Pose and Shape Estimation by Dense Render-and-Compare
- Dual-stream Spatio-Temporal GCN-Transformer Network for 3D Human Pose Estimation
- Anatomy-aware 3D Human Pose Estimation with Bone-based Pose Decomposition
- Learning Pyramid-structured Long-range Dependencies for 3D Human Pose Estimation
- Homogeneous vector bundles and -equivariant convolutional neural networks
- Not All Parts Are Created Equal: 3D Pose Estimation by Modelling Bi-directional Dependencies of Body Parts
- Convex Optimisation for Inverse Kinematics
- Going beyond Free Viewpoint: Creating Animatable Volumetric Video of Human Performances
- 3D Human Pose Estimation using Spatio-Temporal Networks with Explicit Occlusion Training
- GaitVibe+: Enhancing Structural Vibration-based Footstep Localization Using Temporary Cameras for In-home Gait Analysis
- Learning to Reconstruct People in Clothing from a Single RGB Camera
- Shape-Aware Human Pose and Shape Reconstruction Using Multi-View Images
- Single-Network Whole-Body Pose Estimation
- PoseAug: A Differentiable Pose Augmentation Framework for 3D Human Pose Estimation
- Cross-View Tracking for Multi-Human 3D Pose Estimation at over 100 FPS
- FBI-Pose: Towards Bridging the Gap between 2D Images and 3D Human Poses using Forward-or-Backward Information
- A Graph Attention Spatio-temporal Convolutional Network for 3D Human Pose Estimation in Video
- Ego-Pose Estimation and Forecasting as Real-Time PD Control
- Motion Guided 3D Pose Estimation from Videos
- What Face and Body Shapes Can Tell About Height
- Learning 3D Human Dynamics from Video
- Neural Rendering and Reenactment of Human Actor Videos
- Object Activity Scene Description, Construction and Recognition
- LiveCap: Real-time Human Performance Capture from Monocular Video
- Vision-based Estimation of MDS-UPDRS Gait Scores for Assessing Parkinson's Disease Motor Severity
- Human Image Generation: A Comprehensive Survey
- MOVIN: Real-time Motion Capture using a Single LiDAR
- Monocular Real-time Hand Shape and Motion Capture using Multi-modal Data
- GANerated Hands for Real-time 3D Hand Tracking from Monocular RGB
- Temporally Coherent Full 3D Mesh Human Pose Recovery from Monocular Video
- SiCloPe: Silhouette-Based Clothed People
- DeepFuse: An IMU-Aware Network for Real-Time 3D Human Pose Estimation from Multi-View Image
- Category-Level Articulated Object Pose Estimation
- 3D Human Pose Estimation with 2D Marginal Heatmaps
- Learning 3D Human Pose from Structure and Motion
- Conditional Directed Graph Convolution for 3D Human Pose Estimation
- Generalizing Monocular 3D Human Pose Estimation in the Wild
- Perceiving 3D Human-Object Spatial Arrangements from a Single Image in the Wild
- EVOPOSE: A Recursive Transformer For 3D Human Pose Estimation With Kinematic Structure Priors
- Task-Oriented Hand Motion Retargeting for Dexterous Manipulation Imitation
- Human Motion Analysis with Deep Metric Learning
- SMAP: Single-Shot Multi-Person Absolute 3D Pose Estimation
- Imitation Learning for Human Pose Prediction
- Graph and Temporal Convolutional Networks for 3D Multi-person Pose Estimation in Monocular Videos
- MonoClothCap: Towards Temporally Coherent Clothing Capture from Monocular RGB Video
- Single Image Human Proxemics Estimation for Visual Social Distancing
- Monocular Real-Time Volumetric Performance Capture
- NeuralHumanFVV: Real-Time Neural Volumetric Human Performance Rendering using RGB Cameras
- Task-Generic Hierarchical Human Motion Prior using VAEs
- ELMO: Enhanced Real-time LiDAR Motion Capture through Upsampling
- Everybody Is Unique: Towards Unbiased Human Mesh Recovery
- Deformation-aware Unpaired Image Translation for Pose Estimation on Laboratory Animals
- Can Action be Imitated? Learn to Reconstruct and Transfer Human Dynamics from Videos
- StarMap for Category-Agnostic Keypoint and Viewpoint Estimation
- Ordinal Depth Supervision for 3D Human Pose Estimation
- Analysis of Deep-Learning Methods in an ISO/TS 15066-Compliant Human-Robot Safety Framework
- Patch-based 3D Human Pose Refinement
- EventCap: Monocular 3D Capture of High-Speed Human Motions using an Event Camera
- Generative Models for Pose Transfer
- Context Modeling in 3D Human Pose Estimation: A Unified Perspective
- Multi-Scale Networks for 3D Human Pose Estimation with Inference Stage Optimization
- Hierarchical Kinematic Human Mesh Recovery
- Learning the Depths of Moving People by Watching Frozen People
- SimPoE: Simulated Character Control for 3D Human Pose Estimation
- Monocular 3D Multi-Person Pose Estimation by Integrating Top-Down and Bottom-Up Networks
- Self-Supervised Human Depth Estimation from Monocular Videos
- Reconstructing NBA Players
- Real-time Human Finger Pointing Recognition and Estimation for Robot Directives Using a Single Web-Camera
- Structure from Recurrent Motion: From Rigidity to Recurrency
- Lightweight 3D Human Pose Estimation Network Training Using Teacher-Student Learning
- Nonparametric Structure Regularization Machine for 2D Hand Pose Estimation
- 3D Human Pose Estimation Based on 2D-3D Consistency with Synchronized Adversarial Training
- Action Recognition with Spatio-Temporal Visual Attention on Skeleton Image Sequences
- Out of the Box: A combined approach for handling occlusion in Human Pose Estimation
- RobustFusion: Robust Volumetric Performance Reconstruction under Human-object Interactions from Monocular RGBD Stream
- GLAMR: Global Occlusion-Aware Human Mesh Recovery with Dynamic Cameras
- Kinematic-Structure-Preserved Representation for Unsupervised 3D Human Pose Estimation
- LBS Autoencoder: Self-supervised Fitting of Articulated Meshes to Point Clouds
- Single-Shot Multi-Person 3D Pose Estimation From Monocular RGB
- Reconstruction of People in Loose Clothing
- MetaSketch: Wireless Semantic Segmentation by Metamaterial Surfaces
- A Neural Network for Detailed Human Depth Estimation from a Single Image
- Holistic 3D Human and Scene Mesh Estimation from Single View Images
- PCLs: Geometry-aware Neural Reconstruction of 3D Pose with Perspective Crop Layers
- 3D Human Pose Estimation for Free-form Activity Using WiFi Signals
- Computer Vision and Abnormal Patient Gait Assessment a Comparison of Machine Learning Models
- Learning Transferable Kinematic Dictionary for 3D Human Pose and Shape Reconstruction
- PedX: Benchmark Dataset for Metric 3D Pose Estimation of Pedestrians in Complex Urban Intersections
- Deep Reinforcement Learning for Active Human Pose Estimation
- Deep3DPose: Realtime Reconstruction of Arbitrarily Posed Human Bodies from Single RGB Images
- Explicit Spatiotemporal Joint Relation Learning for Tracking Human Pose
- 3D Human Body Reconstruction from a Single Image via Volumetric Regression
- Real-time RGBD-based Extended Body Pose Estimation
- PoseKernelLifter: Metric Lifting of 3D Human Pose using Sound
- Contact and Human Dynamics from Monocular Video
- EgoRenderer: Rendering Human Avatars from Egocentric Camera Images
- Learnable Triangulation for Deep Learning-based 3D Reconstruction of Objects of Arbitrary Topology from Single RGB Images
- Human Body Model Fitting by Learned Gradient Descent
- Soccer on Your Tabletop
- A Tracking System For Baseball Game Reconstruction
- Object Properties Inferring from and Transfer for Human Interaction Motions
- A Deeper Look into DeepCap
- Video Motion Capture from the Part Confidence Maps of Multi-Camera Images by Spatiotemporal Filtering Using the Human Skeletal Model
- Exploring Pose Priors for Human Pose Estimation with Joint Angle Representations