Flowing ConvNets for Human Pose Estimation in Videos
arXiv:1506.02897
Abstract
The objective of this work is human pose estimation in videos, where multiple frames are available. We investigate a ConvNet architecture that is able to benefit from temporal context by combining information across the multiple frames using optical flow. To this end we propose a network architecture with the following novelties: (i) a deeper network than previously investigated for regressing heatmaps; (ii) spatial fusion layers that learn an implicit spatial model; (iii) optical flow is used to align heatmap predictions from neighbouring frames; and (iv) a final parametric pooling layer which learns to combine the aligned heatmaps into a pooled confidence map. We show that this architecture outperforms a number of others, including one that uses optical flow solely at the input layers, one that regresses joint coordinates directly, and one that predicts heatmaps without spatial fusion. The new architecture outperforms the state of the art by a large margin on three video pose estimation datasets, including the very challenging Poses in the Wild dataset, and outperforms other deep methods that don't use a graphical model on the single-image FLIC benchmark (and also Chen & Yuille and Tompson et al. in the high precision region).
ICCV'15
References in corpus (5)
- Two-Stream Convolutional Networks for Action Recognition in Videos
- Conditional Random Fields as Recurrent Neural Networks
- Synthetic Data and Artificial Neural Networks for Natural Scene Text Recognition
- Deep Structured Output Learning for Unconstrained Text Recognition
- MoDeep: A Deep Learning Framework Using Motion Features for Human Pose Estimation
Cited by in corpus (41)
- Human pose estimation via Convolutional Part Heatmap Regression
- Stacked Hourglass Networks for Human Pose Estimation
- Convolutional Pose Machines
- Human Pose Regression by Combining Indirect Part Detection and Contextual Information
- Convolutional Neural Fabrics
- Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields
- Look into Person: Self-supervised Structure-sensitive Learning and A New Benchmark for Human Parsing
- Structured Prediction of 3D Human Pose with Deep Neural Networks
- Large Pose 3D Face Reconstruction from a Single Image via Direct Volumetric CNN Regression
- Sparseness Meets Deepness: 3D Human Pose Estimation from Monocular Video
- Fast and Robust Multi-Person 3D Pose Estimation from Multiple Views
- Going Deeper into First-Person Activity Recognition
- End-to-end Flow Correlation Tracking with Spatial-temporal Attention
- Multi-Scale Structure-Aware Network for Human Pose Estimation
- Learning to Estimate 3D Human Pose and Shape from a Single Color Image
- MonoCap: Monocular Human Motion Capture using a CNN Coupled with a Geometric Prior
- Concurrent Segmentation and Localization for Tracking of Surgical Instruments
- Learning to Fuse 2D and 3D Image Cues for Monocular Body Pose Estimation
- Pose from Action: Unsupervised Learning of Pose Features based on Motion
- Coarse-to-Fine Volumetric Prediction for Single-Image 3D Human Pose
- Thin-Slicing Network: A Deep Structured Model for Pose Estimation in Videos
- Signs in time: Encoding human motion as a temporal image
- End-to-end training of object class detectors for mean average precision
- Reconstructing Vechicles from a Single Image: Shape Priors for Road Scene Understanding
- EgoCap: Egocentric Marker-less Motion Capture with Two Fisheye Cameras
- Real-time Human Pose Estimation from Video with Convolutional Neural Networks
- Learning Human Pose Models from Synthesized Data for Robust RGB-D Action Recognition
- Surveillance Video Parsing with Single Frame Supervision
- PoseTrack: Joint Multi-Person Pose Estimation and Tracking
- Fully-automated deep learning slice-based muscle estimation from CT images for sarcopenia assessment
- Ordinal Depth Supervision for 3D Human Pose Estimation
- Attaining human-level performance with atlas location autocontext for anatomical landmark detection in 3D CT data
- Human Pose Forecasting via Deep Markov Models
- A Supervised Learning Methodology for Real-Time Disguised Face Recognition in the Wild
- Evaluation of Deep Learning based Pose Estimation for Sign Language Recognition
- Beyond Planar Symmetry: Modeling human perception of reflection and rotation symmetries in the wild
- Computer Vision and Abnormal Patient Gait Assessment a Comparison of Machine Learning Models
- ClusterNet: Detecting Small Objects in Large Scenes by Exploiting Spatio-Temporal Information
- Keep it SMPL: Automatic Estimation of 3D Human Pose and Shape from a Single Image
- Single upper limb pose estimation method based on improved stacked hourglass network
- Am I a Baller? Basketball Performance Assessment from First-Person Videos