DeepCut: Joint Subset Partition and Labeling for Multi Person Pose Estimation
arXiv:1511.06645
Abstract
This paper considers the task of articulated human pose estimation of multiple people in real world images. We propose an approach that jointly solves the tasks of detection and pose estimation: it infers the number of persons in a scene, identifies occluded body parts, and disambiguates body parts between people in close proximity of each other. This joint formulation is in contrast to previous strategies, that address the problem by first detecting people and subsequently estimating their body pose. We propose a partitioning and labeling formulation of a set of body-part hypotheses generated with CNN-based part detectors. Our formulation, an instance of an integer linear program, implicitly performs non-maximum suppression on the set of part candidates and groups them to form configurations of body parts respecting geometric and appearance constraints. Experiments on four different datasets demonstrate state-of-the-art results for both single person and multi person pose estimation. Models and code available at http://pose.mpi-inf.mpg.de.
Accepted at IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016)
References in corpus (1)
Cited by in corpus (25)
- Human pose estimation via Convolutional Part Heatmap Regression
- Stacked Hourglass Networks for Human Pose Estimation
- Convolutional Pose Machines
- Pose Invariant Embedding for Deep Person Re-identification
- DeeperCut: A Deeper, Stronger, and Faster Multi-Person Pose Estimation Model
- Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields
- Learning Feature Pyramids for Human Pose Estimation
- Adversarial PoseNet: A Structure-aware Convolutional Network for Human Pose Estimation
- Unite the People: Closing the Loop Between 3D and 2D Human Representations
- Fast and Robust Multi-Person 3D Pose Estimation from Multiple Views
- Human Pose Estimation using Deep Consensus Voting
- Learning to Estimate 3D Human Pose and Shape from a Single Color Image
- Learning to Fuse 2D and 3D Image Cues for Monocular Body Pose Estimation
- Monocular 3D Human Pose Estimation In The Wild Using Improved CNN Supervision
- Thin-Slicing Network: A Deep Structured Model for Pose Estimation in Videos
- Pose2Instance: Harnessing Keypoints for Person Instance Segmentation
- Monocular, One-stage, Regression of Multiple 3D People
- Real-time Human Pose Estimation from Video with Convolutional Neural Networks
- PoseTrack: Joint Multi-Person Pose Estimation and Tracking
- EfficientHRNet: Efficient Scaling for Lightweight High-Resolution Multi-Person Pose Estimation
- Rethinking Pose in 3D: Multi-stage Refinement and Recovery for Markerless Motion Capture
- Ordinal Depth Supervision for 3D Human Pose Estimation
- Coherent Reconstruction of Multiple Humans from a Single Image
- Soccer on Your Tabletop
- Deep Convolutional Poses for Human Interaction Recognition in Monocular Videos