DeeperCut: A Deeper, Stronger, and Faster Multi-Person Pose Estimation Model
arXiv:1605.03170
Abstract
The goal of this paper is to advance the state-of-the-art of articulated pose estimation in scenes with multiple people. To that end we contribute on three fronts. We propose (1) improved body part detectors that generate effective bottom-up proposals for body parts; (2) novel image-conditioned pairwise terms that allow to assemble the proposals into a variable number of consistent body part configurations; and (3) an incremental optimization strategy that explores the search space more efficiently thus leading both to better performance and significant speed-up factors. Evaluation is done on two single-person and two multi-person pose estimation benchmarks. The proposed approach significantly outperforms best known multi-person pose estimation results while demonstrating competitive performance on the task of single person pose estimation. Models and code available at http://pose.mpi-inf.mpg.de
ECCV'16. High-res version at https://www.d2.mpi-inf.mpg.de/sites/default/files/insafutdinov16arxiv.pdf
References in corpus (3)
Cited by in corpus (40)
- VNect: Real-time 3D Human Pose Estimation with a Single RGB Camera
- Human pose estimation via Convolutional Part Heatmap Regression
- Human Pose Regression by Combining Indirect Part Detection and Contextual Information
- Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields
- RMPE: Regional Multi-person Pose Estimation
- Towards human-level performance on automatic pose estimation of infant spontaneous movements
- Multi-Person Pose Estimation with Local Joint-to-Person Associations
- ActionXPose: A Novel 2D Multi-view Pose-based Algorithm for Real-time Human Action Recognition
- Skepxels: Spatio-temporal Image Representation of Human Skeleton Joints for Action Recognition
- ConvNets and ImageNet Beyond Accuracy: Understanding Mistakes and Uncovering Biases
- Integral Human Pose Regression
- SuPer Deep: A Surgical Perception Framework for Robotic Tissue Manipulation using Deep Learning for Feature Extraction
- Joint Multi-Person Pose Estimation and Semantic Part Segmentation
- Multi-Stage HRNet: Multiple Stage High-Resolution Network for Human Pose Estimation
- RePose: Learning Deep Kinematic Priors for Fast Human Pose Estimation
- Real-time Human Pose Estimation from Video with Convolutional Neural Networks
- EgoCap: Egocentric Marker-less Motion Capture with Two Fisheye Cameras
- 4D Association Graph for Realtime Multi-person Motion Capture Using Multiple Video Cameras
- Combining detection and tracking for human pose estimation in videos
- Backbone Can Not be Trained at Once: Rolling Back to Pre-trained Network for Person Re-Identification
- AMIL: Adversarial Multi Instance Learning for Human Pose Estimation
- Multi-scale Aggregation R-CNN for 2D Multi-person Pose Estimation
- EfficientHRNet: Efficient Scaling for Lightweight High-Resolution Multi-Person Pose Estimation
- Toward safe separation distance monitoring from RGB-D sensors in human-robot interaction
- A Multi-view RGB-D Approach for Human Pose Estimation in Operating Rooms
- Collaborative Attention Network for Person Re-identification
- Exploiting skeletal structure in computer vision annotation with Benders decomposition
- Multi-Person Pose Estimation via Column Generation
- Bi-directional Graph Structure Information Model for Multi-Person Pose Estimation
- Improving Temporal Interpolation of Head and Body Pose using Gaussian Process Regression in a Matrix Completion Setting
- Efficient Pose and Cell Segmentation using Column Generation
- PedX: Benchmark Dataset for Metric 3D Pose Estimation of Pedestrians in Complex Urban Intersections
- Single-Shot Multi-Person 3D Pose Estimation From Monocular RGB
- Flow-Partitionable Signed Graphs
- Inter-Homines: Distance-Based Risk Estimation for Human Safety
- Simple Multi-Resolution Representation Learning for Human Pose Estimation
- Bottom-up Pose Estimation of Multiple Person with Bounding Box Constraint
- Motion Style Extraction Based on Sparse Coding Decomposition
- A Global to Local Double Embedding Method for Multi-person Pose Estimation
- An Empirical Study towards Understanding How Deep Convolutional Nets Recognize Falls