Convolutional Pose Machines
arXiv:1602.00134
Abstract
Pose Machines provide a sequential prediction framework for learning rich implicit spatial models. In this work we show a systematic design for how convolutional networks can be incorporated into the pose machine framework for learning image features and image-dependent spatial models for the task of pose estimation. The contribution of this paper is to implicitly model long-range dependencies between variables in structured prediction tasks such as articulated pose estimation. We achieve this by designing a sequential architecture composed of convolutional networks that directly operate on belief maps from previous stages, producing increasingly refined estimates for part locations, without the need for explicit graphical model-style inference. Our approach addresses the characteristic difficulty of vanishing gradients during training by providing a natural learning objective function that enforces intermediate supervision, thereby replenishing back-propagated gradients and conditioning the learning procedure. We demonstrate state-of-the-art performance and outperform competing methods on standard benchmarks including the MPII, LSP, and FLIC datasets.
camera ready
References in corpus (4)
Cited by in corpus (116)
- Human pose estimation via Convolutional Part Heatmap Regression
- Stacked Hourglass Networks for Human Pose Estimation
- GLAD: Global-Local-Alignment Descriptor for Pedestrian Retrieval
- Human Pose Regression by Combining Indirect Part Detection and Contextual Information
- Deep Pictorial Gaze Estimation
- MoCap-guided Data Augmentation for 3D Pose Estimation in the Wild
- Numerical Coordinate Regression with Convolutional Neural Networks
- Pose Invariant Embedding for Deep Person Re-identification
- Deeply-Learned Part-Aligned Representations for Person Re-Identification
- AI Challenger : A Large-scale Dataset for Going Deeper in Image Understanding
- Self-supervised Learning of Motion Capture
- DeeperCut: A Deeper, Stronger, and Faster Multi-Person Pose Estimation Model
- Rethinking Feature Discrimination and Polymerization for Large-scale Recognition
- A simple yet effective baseline for 3d human pose estimation
- PixelNet: Representation of the pixels, by the pixels, and for the pixels
- Towards Accurate Multi-person Pose Estimation in the Wild
- Towards 3D Human Pose Estimation in the Wild: a Weakly-supervised Approach
- Synthesizing Training Images for Boosting Human 3D Pose Estimation
- BodyNet: Volumetric Inference of 3D Human Body Shapes
- Learning Feature Pyramids for Human Pose Estimation
- RMPE: Regional Multi-person Pose Estimation
- Adversarial PoseNet: A Structure-aware Convolutional Network for Human Pose Estimation
- Detecting and Recognizing Human-Object Interactions
- Style Aggregated Network for Facial Landmark Detection
- Unite the People: Closing the Loop Between 3D and 2D Human Representations
- Jointly Optimize Data Augmentation and Network Training: Adversarial Data Augmentation in Human Pose Estimation
- Real-Time Human Pose Estimation on a Smart Walker using Convolutional Neural Networks
- Fast and Robust Multi-Person 3D Pose Estimation from Multiple Views
- 3D Human Pose Estimation in the Wild by Adversarial Learning
- Supervision-by-Registration: An Unsupervised Approach to Improve the Precision of Facial Landmark Detectors
- Action Machine: Rethinking Action Recognition in Trimmed Videos
- Optical Flow Guided Feature: A Fast and Robust Motion Representation for Video Action Recognition
- Iterative Visual Reasoning Beyond Convolutions
- Video Based Reconstruction of 3D People Models
- Monocular Total Capture: Posing Face, Body, and Hands in the Wild
- Attention-based Context Aggregation Network for Monocular Depth Estimation
- Estimating 6D Pose From Localizing Designated Surface Keypoints
- Detailed, accurate, human shape estimation from clothed 3D scan sequences
- Multi-Person Pose Estimation with Local Joint-to-Person Associations
- A Multi-Stream Convolutional Neural Network Framework for Group Activity Recognition
- First-Person Hand Action Benchmark with RGB-D Videos and 3D Hand Pose Annotations
- Learning Discriminative Motion Features Through Detection
- Unsupervised Person Image Synthesis in Arbitrary Poses
- Integral Human Pose Regression
- Holistic Planimetric prediction to Local Volumetric prediction for 3D Human Pose Estimation
- Coarse-to-Fine Volumetric Prediction for Single-Image 3D Human Pose
- RGB-based 3D Hand Pose Estimation via Privileged Learning with Depth Images
- EraseReLU: A Simple Way to Ease the Training of Deep Convolution Neural Networks
- Let's Dance: Learning From Online Dance Videos
- 3D Human Pose Estimation from a Single Image via Distance Matrix Regression
- Thin-Slicing Network: A Deep Structured Model for Pose Estimation in Videos
- ApolloCar3D: A Large 3D Car Instance Understanding Benchmark for Autonomous Driving
- 3D Human Pose Estimation = 2D Pose Estimation + Matching
- Quantized Densely Connected U-Nets for Efficient Landmark Localization
- FBI-Pose: Towards Bridging the Gap between 2D Images and 3D Human Poses using Forward-or-Backward Information
- LightTrack: A Generic Framework for Online Top-Down Human Pose Tracking
- Improving Multi-Person Pose Estimation using Label Correction
- Generative Partition Networks for Multi-Person Pose Estimation
- Dual Path Networks for Multi-Person Human Pose Estimation
- Visual Compiler: Synthesizing a Scene-Specific Pedestrian Detector and Pose Estimator
- Learning to Refine Human Pose Estimation
- Learning Detailed Face Reconstruction from a Single Image
- PoseTrack: Joint Multi-Person Pose Estimation and Tracking
- EgoCap: Egocentric Marker-less Motion Capture with Two Fisheye Cameras
- Real-time Human Pose Estimation from Video with Convolutional Neural Networks
- Phonology Recognition in American Sign Language
- Deep Dose Plugin Towards Real-time Monte Carlo Dose Calculation Through a Deep Learning based Denoising Algorithm
- DeepSkeleton: Skeleton Map for 3D Human Pose Regression
- Photo Wake-Up: 3D Character Animation from a Single Photo
- Attributes-aided Part Detection and Refinement for Person Re-identification
- Skeleton-based Gesture Recognition Using Several Fully Connected Layers with Path Signature Features and Temporal Transformer Module
- A Multi-view RGB-D Approach for Human Pose Estimation in Operating Rooms
- EfficientHRNet: Efficient Scaling for Lightweight High-Resolution Multi-Person Pose Estimation
- Deep Kinematic Pose Regression
- Dense 3D Regression for Hand Pose Estimation
- DeepFlux for Skeletons in the Wild
- Turbo Learning Framework for Human-Object Interactions Recognition and Human Pose Estimation
- Real Time Fine-Grained Categorization with Accuracy and Interpretability
- StarMap for Category-Agnostic Keypoint and Viewpoint Estimation
- Rethinking Pose in 3D: Multi-stage Refinement and Recovery for Markerless Motion Capture
- Landmark Detection in Low Resolution Faces with Semi-Supervised Learning
- Train Your Data Processor: Distribution-Aware and Error-Compensation Coordinate Decoding for Human Pose Estimation
- Finger Grip Force Estimation from Video using Two Stream Approach
- When Vehicles See Pedestrians with Phones:A Multi-Cue Framework for Recognizing Phone-based Activities of Pedestrians
- Ordinal Depth Supervision for 3D Human Pose Estimation
- Post-Data Augmentation to Improve Deep Pose Estimation of Extreme and Wild Motions
- Disjoint Multi-task Learning between Heterogeneous Human-centric Tasks
- Human Pose Forecasting via Deep Markov Models
- Higher-order Pooling of CNN Features via Kernel Linearization for Action Recognition
- Improving Temporal Interpolation of Head and Body Pose using Gaussian Process Regression in a Matrix Completion Setting
- Unstructured Road Vanishing Point Detection Using the Convolutional Neural Network and Heatmap Regression
- View Invariant 3D Human Pose Estimation
- Bi-directional Graph Structure Information Model for Multi-Person Pose Estimation
- A Video-Based Method for Objectively Rating Ataxia
- Recurrent 3D Pose Sequence Machines
- Computer Vision and Abnormal Patient Gait Assessment a Comparison of Machine Learning Models
- Perceive Where to Focus: Learning Visibility-aware Part-level Features for Partial Person Re-identification
- Deformable Part Networks
- Keep it SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image
- Beyond Planar Symmetry: Modeling human perception of reflection and rotation symmetries in the wild
- 3D Human Pose Estimation Using Convolutional Neural Networks with 2D Pose Information
- Generalized Coarse-to-Fine Visual Recognition with Progressive Training
- Efficient Human Pose Estimation by Learning Deeply Aggregated Representations
- Deep Reinforcement Learning for Active Human Pose Estimation
- Human Pose Estimation using Global and Local Normalization
- Inner Space Preserving Generative Pose Machine
- Hand-tremor frequency estimation in videos
- Simple Multi-Resolution Representation Learning for Human Pose Estimation
- HANDS18: Methods, Techniques and Applications for Hand Observation
- Soccer on Your Tabletop
- Learning to Estimate Pose by Watching Videos
- Joint Voxel and Coordinate Regression for Accurate 3D Facial Landmark Localization
- Deep Convolutional Poses for Human Interaction Recognition in Monocular Videos
- Linear Span Network for Object Skeleton Detection
- ActiveNet: A computer-vision based approach to determine lethargy
- Bottom-up Pose Estimation of Multiple Person with Bounding Box Constraint