Learning Human Pose Estimation Features with Convolutional Networks
arXiv:1312.7302
Abstract
This paper introduces a new architecture for human pose estimation using a multi- layer convolutional network architecture and a modified learning technique that learns low-level features and higher-level weak spatial models. Unconstrained human pose estimation is one of the hardest problems in computer vision, and our new architecture and learning schema shows significant improvement over the current state-of-the-art results. The main contribution of this paper is showing, for the first time, that a specific variation of deep learning is able to outperform all existing traditional architectures on this task. The paper also discusses several lessons learned while researching alternatives, most notably, that it is possible to learn strong low-level feature detectors on features that might even just cover a few pixels in the image. Higher-level spatial models improve somewhat the overall result, but to a much lesser extent then expected. Many researchers previously argued that the kinematic structure and top-down information is crucial for this domain, but with our purely bottom up, and weak spatial model, we could improve other more complicated architectures that currently produce the best results. This mirrors what many other researchers, like those in the speech recognition, object recognition, and other domains have experienced.
References in corpus (1)
Cited by in corpus (42)
- Joint Training of a Convolutional Network and a Graphical Model for Human Pose Estimation
- Lifting from the Deep: Convolutional 3D Pose Estimation from a Single Image
- Monocular Human Pose Estimation: A Survey of Deep Learning-based Methods
- Recent Advances in Convolutional Neural Networks
- Do Convnets Learn Correspondence?
- Learning Spatiotemporal Features with 3D Convolutional Networks
- MoVi: A Large Multipurpose Motion and Video Dataset
- Combining Local Appearance and Holistic View: Dual-Source Deep Neural Networks for Human Pose Estimation
- Flowing ConvNets for Human Pose Estimation in Videos
- Towards Accurate Multi-person Pose Estimation in the Wild
- RMPE: Regional Multi-person Pose Estimation
- Robust 3D Hand Pose Estimation in Single Depth Images: from Single-View CNN to Multi-View CNNs
- Structured Prediction of 3D Human Pose with Deep Neural Networks
- Sparseness Meets Deepness: 3D Human Pose Estimation from Monocular Video
- Part Detector Discovery in Deep Convolutional Neural Networks
- MoDeep: A Deep Learning Framework Using Motion Features for Human Pose Estimation
- Maximum-Margin Structured Learning with Deep Networks for 3D Human Pose Estimation
- Video-based Human Action Recognition using Deep Learning: A Review
- Estimating 6D Pose From Localizing Designated Surface Keypoints
- Matching-CNN Meets KNN: Quasi-Parametric Human Parsing
- Ego3DPose: Capturing 3D Cues from Binocular Egocentric Views
- MonoCap: Monocular Human Motion Capture using a CNN Coupled with a Geometric Prior
- Learning to Fuse 2D and 3D Image Cues for Monocular Body Pose Estimation
- Training a Feedback Loop for Hand Pose Estimation
- Learning Fine-grained Features via a CNN Tree for Large-scale Classification
- Learning 3D Human Pose from Structure and Motion
- Real-time Human Pose Estimation from Video with Convolutional Neural Networks
- EgoCap: Egocentric Marker-less Motion Capture with Two Fisheye Cameras
- On the Robustness of Human Pose Estimation
- EfficientHRNet: Efficient Scaling for Lightweight High-Resolution Multi-Person Pose Estimation
- A Multi-view RGB-D Approach for Human Pose Estimation in Operating Rooms
- Multi-person 3D pose estimation from unlabelled data
- SimPose: Effectively Learning DensePose and Surface Normals of People from Simulated Data
- Single Person Pose Estimation: A Survey
- Bi-directional Graph Structure Information Model for Multi-Person Pose Estimation
- Lightweight 3D Human Pose Estimation Network Training Using Teacher-Student Learning
- Evaluation of Deep Learning based Pose Estimation for Sign Language Recognition
- Enhanced Mixtures of Part Model for Human Pose Estimation
- Multi-Person Pose Estimation with Enhanced Feature Aggregation and Selection
- Deep Spatio-temporal Manifold Network for Action Recognition
- Towards Viewpoint Invariant 3D Human Pose Estimation
- Learning to Estimate Pose by Watching Videos