R-CNNs for Pose Estimation and Action Detection
arXiv:1406.5212
Abstract
We present convolutional neural networks for the tasks of keypoint (pose) prediction and action classification of people in unconstrained images. Our approach involves training an R-CNN detector with loss functions depending on the task being tackled. We evaluate our method on the challenging PASCAL VOC dataset and compare it to previous leading approaches. Our method gives state-of-the-art results for keypoint and action prediction. Additionally, we introduce a new dataset for action detection, the task of simultaneously localizing people and classifying their actions, and present results using our approach.
Cited by in corpus (14)
- Monocular Human Pose Estimation: A Survey of Deep Learning-based Methods
- Dynamic Feature Integration for Simultaneous Detection of Salient Object, Edge and Skeleton
- Video-based Human Action Recognition using Deep Learning: A Review
- Hypercolumns for Object Segmentation and Fine-grained Localization
- Group Ensemble: Learning an Ensemble of ConvNets in a single ConvNet
- PaStaNet: Toward Human Activity Knowledge Engine
- Intention Recognition of Pedestrians and Cyclists by 2D Pose Estimation
- Single Image Action Recognition by Predicting Space-Time Saliency
- Cross-Domain Self-supervised Multi-task Feature Learning using Synthetic Imagery
- MPM: Joint Representation of Motion and Position Map for Cell Tracking
- Autonomous Navigation in Dynamic Environments: Deep Learning-Based Approach
- Learning Action Concept Trees and Semantic Alignment Networks from Image-Description Data
- Understanding and Predicting The Attractiveness of Human Action Shot
- Nuisance-Label Supervision: Robustness Improvement by Free Labels