Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields
arXiv:1611.08050
Abstract
We present an approach to efficiently detect the 2D pose of multiple people in an image. The approach uses a nonparametric representation, which we refer to as Part Affinity Fields (PAFs), to learn to associate body parts with individuals in the image. The architecture encodes global context, allowing a greedy bottom-up parsing step that maintains high accuracy while achieving realtime performance, irrespective of the number of people in the image. The architecture is designed to jointly learn part locations and their association via two branches of the same sequential prediction process. Our method placed first in the inaugural COCO 2016 keypoints challenge, and significantly exceeds the previous state-of-the-art result on the MPII Multi-Person benchmark, both in performance and efficiency.
Accepted as CVPR 2017 Oral. Video result: https://youtu.be/pW6nZXeWlGM
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Human pose estimation via Convolutional Part Heatmap Regression
- Stacked Hourglass Networks for Human Pose Estimation
- DeeperCut: A Deeper, Stronger, and Faster Multi-Person Pose Estimation Model
- Efficient Object Localization Using Convolutional Networks
- Multi-Person Pose Estimation with Local Joint-to-Person Associations
Cited by in corpus (32)
- VNect: Real-time 3D Human Pose Estimation with a Single RGB Camera
- Object Detection in 20 Years: A Survey
- Revisiting Unreasonable Effectiveness of Data in Deep Learning Era
- PixelNet: Representation of the pixels, by the pixels, and for the pixels
- Towards Accurate Multi-person Pose Estimation in the Wild
- Scanner: Efficient Video Analysis at Scale
- RMPE: Regional Multi-person Pose Estimation
- Deep Cocktail Network: Multi-source Unsupervised Domain Adaptation with Category Shift
- Pose is all you need: The pose only group activity recognition system (POGARS)
- Compact Real-time avoidance on a Humanoid Robot for Human-robot Interaction
- Point Linking Network for Object Detection
- Integral Human Pose Regression
- Dual Path Networks for Multi-Person Human Pose Estimation
- GODS: Generalized One-class Discriminative Subspaces for Anomaly Detection
- Learning to Engage with Interactive Systems: A Field Study on Deep Reinforcement Learning in a Public Museum
- Towards Disentangled Representations for Human Retargeting by Multi-view Learning
- TRB: A Novel Triplet Representation for Understanding 2D Human Body
- A Proposed Set of Communicative Gestures for Human Robot Interaction and an RGB Image-based Gesture Recognizer Implemented in ROS
- When Vehicles See Pedestrians with Phones:A Multi-Cue Framework for Recognizing Phone-based Activities of Pedestrians
- Disjoint Multi-task Learning between Heterogeneous Human-centric Tasks
- Being the center of attention: A Person-Context CNN framework for Personality Recognition
- Effects of Interruptibility-Aware Robot Behavior
- Hierarchical Model for Long-term Video Prediction
- Action Recognition with Spatio-Temporal Visual Attention on Skeleton Image Sequences
- Human Pose Forecasting via Deep Markov Models
- Object Detection via Aspect Ratio and Context Aware Region-based Convolutional Networks
- Unifying Identification and Context Learning for Person Recognition
- Coupled Recurrent Network (CRN)
- Skeletal Data Matching and Merging from Multiple RGB-D Sensors for Room-Scale Human Behaviour Tracking
- Bottom-up Pose Estimation of Multiple Person with Bounding Box Constraint
- Multi-Scale Supervised Network for Human Pose Estimation
- Learning Conditional Random Fields with Augmented Observations for Partially Observed Action Recognition