Stacked Hourglass Networks for Human Pose Estimation
arXiv:1603.06937
Abstract
This work introduces a novel convolutional network architecture for the task of human pose estimation. Features are processed across all scales and consolidated to best capture the various spatial relationships associated with the body. We show how repeated bottom-up, top-down processing used in conjunction with intermediate supervision is critical to improving the performance of the network. We refer to the architecture as a "stacked hourglass" network based on the successive steps of pooling and upsampling that are done to produce a final set of predictions. State-of-the-art results are achieved on the FLIC and MPII benchmarks outcompeting all recent methods.
References in corpus (6)
- Learning Deconvolution Network for Semantic Segmentation
- Indoor Semantic Segmentation using depth information
- Convolutional Pose Machines
- Weakly-supervised Disentangling with Recurrent Transformations for 3D View Synthesis
- Combining Local Appearance and Holistic View: Dual-Source Deep Neural Networks for Human Pose Estimation
- MoDeep: A Deep Learning Framework Using Motion Features for Human Pose Estimation
Cited by in corpus (212)
- Encoder-Decoder with Atrous Separable Convolution for Semantic Image Segmentation
- VNect: Real-time 3D Human Pose Estimation with a Single RGB Camera
- Pose Guided Person Image Generation
- Human pose estimation via Convolutional Part Heatmap Regression
- EV-FlowNet: Self-Supervised Optical Flow Estimation for Event-based Cameras
- Feature Pyramid Networks for Object Detection
- Residual Attention Network for Image Classification
- CalibNet: Geometrically Supervised Extrinsic Calibration using 3D Spatial Transformer Networks
- Pyramid Attention Network for Semantic Segmentation
- Deep Pictorial Gaze Estimation
- Dual Attention Network for Scene Segmentation
- Pixels to Graphs by Associative Embedding
- Numerical Coordinate Regression with Convolutional Neural Networks
- Pose Invariant Embedding for Deep Person Re-identification
- Learning to Generate Long-term Future via Hierarchical Prediction
- AI Challenger : A Large-scale Dataset for Going Deeper in Image Understanding
- GeoNet: Unsupervised Learning of Dense Depth, Optical Flow and Camera Pose
- Simple Baselines for Human Pose Estimation and Tracking
- Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields
- Single-Image Depth Perception in the Wild
- A simple yet effective baseline for 3d human pose estimation
- HydraPlus-Net: Attentive Deep Features for Pedestrian Analysis
- Stacked Deconvolutional Network for Semantic Segmentation
- Towards Accurate Multi-person Pose Estimation in the Wild
- Towards 3D Human Pose Estimation in the Wild: a Weakly-supervised Approach
- Knowledge Distillation in Generations: More Tolerant Teachers Educate Better Students
- Pyramid Stereo Matching Network
- BodyNet: Volumetric Inference of 3D Human Body Shapes
- OriNet: A Fully Convolutional Network for 3D Human Pose Estimation
- MDSSD: Multi-scale Deconvolutional Single Shot Detector for Small Objects
- Learning Feature Pyramids for Human Pose Estimation
- Adversarial PoseNet: A Structure-aware Convolutional Network for Human Pose Estimation
- RMPE: Regional Multi-person Pose Estimation
- TFPose: Direct Human Pose Estimation with Transformers
- Deciding How to Decide: Dynamic Routing in Artificial Neural Networks
- Style Aggregated Network for Facial Landmark Detection
- A Point Set Generation Network for 3D Object Reconstruction from a Single Image
- Adversarial Examples in Modern Machine Learning: A Review
- A Survey on Deep Learning Methods for Robot Vision
- Towards human-level performance on automatic pose estimation of infant spontaneous movements
- Unsupervised Adversarial Learning of 3D Human Pose from 2D Joint Locations
- Large Pose 3D Face Reconstruction from a Single Image via Direct Volumetric CNN Regression
- DAIS: Automatic Channel Pruning via Differentiable Annealing Indicator Search
- Unite the People: Closing the Loop Between 3D and 2D Human Representations
- MonoPerfCap: Human Performance Capture from Monocular Video
- Jointly Optimize Data Augmentation and Network Training: Adversarial Data Augmentation in Human Pose Estimation
- Transformation-Grounded Image Generation Network for Novel 3D View Synthesis
- Real-Time Human Pose Estimation on a Smart Walker using Convolutional Neural Networks
- V2V-PoseNet: Voxel-to-Voxel Prediction Network for Accurate 3D Hand and Human Pose Estimation from a Single Depth Map
- FSRNet: End-to-End Learning Face Super-Resolution with Facial Priors
- Hallucinated-IQA: No-Reference Image Quality Assessment via Adversarial Learning
- Fast and Robust Multi-Person 3D Pose Estimation from Multiple Views
- 3D Human Pose Estimation in the Wild by Adversarial Learning
- Real-time Convolutional Networks for Depth-based Human Pose Estimation
- Stacked U-Nets: A No-Frills Approach to Natural Image Segmentation
- Learning Pose Grammar to Encode Human Body Configuration for 3D Pose Estimation
- Polar Transformer Networks
- Supervision-by-Registration: An Unsupervised Approach to Improve the Precision of Facial Landmark Detectors
- Pose is all you need: The pose only group activity recognition system (POGARS)
- Action Machine: Rethinking Action Recognition in Trimmed Videos
- Optical Flow Guided Feature: A Fast and Robust Motion Representation for Video Action Recognition
- 3D Human Pose Estimation with Relational Networks
- Joint Flow: Temporal Flow Fields for Multi Person Tracking
- Look at Boundary: A Boundary-Aware Face Alignment Algorithm
- Iterative Visual Reasoning Beyond Convolutions
- Recurrent CNN for 3D Gaze Estimation using Appearance and Shape Cues
- Pose-Invariant Face Alignment with a Single CNN
- Zoom Out-and-In Network with Recursive Training for Object Proposal
- Deep Learning For Face Recognition: A Critical Analysis
- Monocular Total Capture: Posing Face, Body, and Hands in the Wild
- Estimating 6D Pose From Localizing Designated Surface Keypoints
- Attention-based Context Aggregation Network for Monocular Depth Estimation
- Multi-Scale Structure-Aware Network for Human Pose Estimation
- Automatic Pixelwise Object Labeling for Aerial Imagery Using Stacked U-Nets
- Detailed, accurate, human shape estimation from clothed 3D scan sequences
- Learning to Estimate 3D Human Pose and Shape from a Single Color Image
- Multi-Person Pose Estimation with Local Joint-to-Person Associations
- Human Pose and Path Estimation from Aerial Video using Dynamic Classifier Selection
- Audio query-based music source separation
- Dynamic Computational Time for Visual Attention
- UAV-Human: A Large Benchmark for Human Behavior Understanding with Unmanned Aerial Vehicles
- Beyond the Pixel-Wise Loss for Topology-Aware Delineation
- Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes
- Mono3D++: Monocular 3D Vehicle Detection with Two-Scale 3D Hypotheses and Task Priors
- Dense Transformer Networks
- Learning to Fuse 2D and 3D Image Cues for Monocular Body Pose Estimation
- GraFormer: Graph Convolution Transformer for 3D Pose Estimation
- Integral Human Pose Regression
- Monocular 3D Human Pose Estimation In The Wild Using Improved CNN Supervision
- Holistic Planimetric prediction to Local Volumetric prediction for 3D Human Pose Estimation
- Coarse-to-Fine Volumetric Prediction for Single-Image 3D Human Pose
- Exchangeable deep neural networks for set-to-set matching and learning
- LineNet: a Zoomable CNN for Crowdsourced High Definition Maps Modeling in Urban Environments
- M2Det: A Single-Shot Object Detector based on Multi-Level Feature Pyramid Network
- Thin-Slicing Network: A Deep Structured Model for Pose Estimation in Videos
- Pose2Seg: Detection Free Human Instance Segmentation
- Iterative Deep Learning for Road Topology Extraction
- FBI-Pose: Towards Bridging the Gap between 2D Images and 3D Human Poses using Forward-or-Backward Information
- Quantized Densely Connected U-Nets for Efficient Landmark Localization
- Towards Robust RGB-D Human Mesh Recovery
- Iterative Deep Learning for Network Topology Extraction
- LightTrack: A Generic Framework for Online Top-Down Human Pose Tracking
- FAB: A Robust Facial Landmark Detection Framework for Motion-Blurred Videos
- Weakly-Supervised 3D Pose Estimation from a Single Image using Multi-View Consistency
- Improving Multi-Person Pose Estimation using Label Correction
- A Learning-based Framework for Hybrid Depth-from-Defocus and Stereo Matching
- Generative Partition Networks for Multi-Person Pose Estimation
- Dual Path Networks for Multi-Person Human Pose Estimation
- Holistic 3D Scene Parsing and Reconstruction from a Single RGB Image
- Joint Multi-Person Pose Estimation and Semantic Part Segmentation
- DRPose3D: Depth Ranking in 3D Human Pose Estimation
- RePose: Learning Deep Kinematic Priors for Fast Human Pose Estimation
- Learning to Refine Human Pose Estimation
- Visual Compiler: Synthesizing a Scene-Specific Pedestrian Detector and Pose Estimator
- CU-Net: Coupled U-Nets
- The Devil is in the Decoder: Classification, Regression and GANs
- Visualized Insights into the Optimization Landscape of Fully Convolutional Networks
- AnchorFace: An Anchor-based Facial Landmark Detector Across Large Poses
- Learning to Forecast and Refine Residual Motion for Image-to-Video Generation
- InverseRenderNet: Learning single image inverse rendering
- Model-based 3D Hand Reconstruction via Self-Supervised Learning
- Accelerating Deep Neural Networks with Spatial Bottleneck Modules
- Real-time Human Pose Estimation from Video with Convolutional Neural Networks
- Glance and Gaze: Inferring Action-aware Points for One-Stage Human-Object Interaction Detection
- Adaloss: Adaptive Loss Function for Landmark Localization
- PoseTrack: Joint Multi-Person Pose Estimation and Tracking
- EgoCap: Egocentric Marker-less Motion Capture with Two Fisheye Cameras
- 3D Human Pose Estimation with 2D Marginal Heatmaps
- Pose-Based Two-Stream Relational Networks for Action Recognition in Videos
- Weakly and Semi Supervised Human Body Part Parsing via Pose-Guided Knowledge Transfer
- Learning 3D Human Pose from Structure and Motion
- BRULÈ: Barycenter-Regularized Unsupervised Landmark Extraction
- End-to-End Deep Kronecker-Product Matching for Person Re-identification
- AOGNets: Compositional Grammatical Architectures for Deep Learning
- Zoom Out-and-In Network with Map Attention Decision for Region Proposal and Object Detection
- DeepSkeleton: Skeleton Map for 3D Human Pose Regression
- Geometry-Aware Face Completion and Editing
- Human Motion Analysis with Deep Metric Learning
- EfficientHRNet: Efficient Scaling for Lightweight High-Resolution Multi-Person Pose Estimation
- Matrix and tensor decompositions for training binary neural networks
- 3D Pose Detection in Videos: Focusing on Occlusion
- Deep Kinematic Pose Regression
- A Multi-view RGB-D Approach for Human Pose Estimation in Operating Rooms
- Towards Highly Accurate and Stable Face Alignment for High-Resolution Videos
- A Fast and Accurate System for Face Detection, Identification, and Verification
- StarMap for Category-Agnostic Keypoint and Viewpoint Estimation
- Dense 3D Regression for Hand Pose Estimation
- Deformation-aware Unpaired Image Translation for Pose Estimation on Laboratory Animals
- Deep Learned Frame Prediction for Video Compression
- Part-Aware Measurement for Robust Multi-View Multi-Human 3D Pose Estimation and Tracking
- Train Your Data Processor: Distribution-Aware and Error-Compensation Coordinate Decoding for Human Pose Estimation
- An End-to-end Framework for Unconstrained Monocular 3D Hand Pose Estimation
- Pedestrian Detection with Autoregressive Network Phases
- Ordinal Depth Supervision for 3D Human Pose Estimation
- Rethinking Pose in 3D: Multi-stage Refinement and Recovery for Markerless Motion Capture
- Chained Predictions Using Convolutional Neural Networks
- Vehicle Re-Identification in Context
- Hierarchical Kinematic Human Mesh Recovery
- Looking for change? Roll the Dice and demand Attention
- Bi-directional Graph Structure Information Model for Multi-Person Pose Estimation
- SRH-Net: Stacked Recurrent Hourglass Network for Stereo Matching
- Vehicle Reconstruction and Texture Estimation Using Deep Implicit Semantic Template Mapping
- Deep Consensus Learning
- Correlating Edge, Pose with Parsing
- Spine Landmark Localization with combining of Heatmap Regression and Direct Coordinate Regression
- Key Frame Proposal Network for Efficient Pose Estimation in Videos
- Toward Marker-free 3D Pose Estimation in Lifting: A Deep Multi-view Solution
- Action Recognition with Spatio-Temporal Visual Attention on Skeleton Image Sequences
- Lightweight 3D Human Pose Estimation Network Training Using Teacher-Student Learning
- View Invariant 3D Human Pose Estimation
- Disentangling Pose from Appearance in Monochrome Hand Images
- High Frequency Residual Learning for Multi-Scale Image Classification
- Hierarchical Back Projection Network for Image Super-Resolution
- Human Pose Forecasting via Deep Markov Models
- Hierarchical Model for Long-term Video Prediction
- Automated Identification of Trampoline Skills Using Computer Vision Extracted Pose Estimation
- Detect, Replace, Refine: Deep Structured Prediction For Pixel Wise Labeling
- ULSD: Unified Line Segment Detection across Pinhole, Fisheye, and Spherical Cameras
- Convolutional Point-set Representation: A Convolutional Bridge Between a Densely Annotated Image and 3D Face Alignment
- Interactive Text2Pickup Network for Natural Language based Human-Robot Collaboration
- Pose estimator and tracker using temporal flow maps for limbs
- Atypical Facial Landmark Localisation with Stacked Hourglass Networks: A Study on 3D Facial Modelling for Medical Diagnosis
- Sharpen Focus: Learning with Attention Separability and Consistency
- Adversarial 3D Human Pose Estimation via Multimodal Depth Supervision
- Human Pose Estimation using Global and Local Normalization
- Motion deblurring of faces
- Computer Vision and Abnormal Patient Gait Assessment a Comparison of Machine Learning Models
- KPNet: Towards Minimal Face Detector
- Human Action Adverb Recognition: ADHA Dataset and A Three-Stream Hybrid Model
- Learning Human Poses from Actions
- PC-HMR: Pose Calibration for 3D Human Mesh Recovery from 2D Images/Videos
- Shape from Shading through Shape Evolution
- Single-Shot Multi-Person 3D Pose Estimation From Monocular RGB
- SyDog: A Synthetic Dog Dataset for Improved 2D Pose Estimation
- Unposed: Unsupervised Pose Estimation based Product Image Recommendations
- ADNet: Leveraging Error-Bias Towards Normal Direction in Face Alignment
- Learning to Predict Diverse Human Motions from a Single Image via Mixture Density Networks
- Perceive Where to Focus: Learning Visibility-aware Part-level Features for Partial Person Re-identification
- Multi-Scale Spatially-Asymmetric Recalibration for Image Classification
- Joint Voxel and Coordinate Regression for Accurate 3D Facial Landmark Localization
- Human Recognition Using Face in Computed Tomography
- Bottom-up Pose Estimation of Multiple Person with Bounding Box Constraint
- Overcoming the Domain Gap in Contrastive Learning of Neural Action Representations
- Improving drone localisation around wind turbines using monocular model-based tracking
- Think about boundary: Fusing multi-level boundary information for landmark heatmap regression
- Soccer on Your Tabletop
- Affinity Derivation and Graph Merge for Instance Segmentation
- A Global to Local Double Embedding Method for Multi-person Pose Estimation
- Inner Space Preserving Generative Pose Machine
- IC-Network: Efficient Structure for Convolutional Neural Networks
- Robust 3D Self-portraits in Seconds
- Layout-Graph Reasoning for Fashion Landmark Detection