PANDA: Pose Aligned Networks for Deep Attribute Modeling
arXiv:1311.5591
Abstract
We propose a method for inferring human attributes (such as gender, hair style, clothes style, expression, action) from images of people under large variation of viewpoint, pose, appearance, articulation and occlusion. Convolutional Neural Nets (CNN) have been shown to perform very well on large scale object recognition problems. In the context of attribute classification, however, the signal is often subtle and it may cover only a small part of the image, while the image is dominated by the effects of pose and viewpoint. Discounting for pose variation would require training on very large labeled datasets which are not presently available. Part-based models, such as poselets and DPM have been shown to perform well for this problem but they are limited by shallow low-level features. We propose a new method which combines part-based models and deep learning by training pose-normalized CNNs. We show substantial improvement vs. state-of-the-art methods on challenging attribute classification tasks in unconstrained settings. Experiments confirm that our method outperforms both the best part-based methods on this problem and conventional CNNs trained on the full bounding box of the person.
8 pages
References in corpus (2)
Cited by in corpus (18)
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Towards Deep Neural Network Architectures Robust to Adversarial Examples
- Learning Spatiotemporal Features with 3D Convolutional Networks
- Deep Learning Face Attributes in the Wild
- HydraPlus-Net: Attentive Deep Features for Pedestrian Analysis
- Co-training for Demographic Classification Using Deep Learning from Label Proportions
- Action Recognition with Image Based CNN Features
- ReconNet: Non-Iterative Reconstruction of Images from Compressively Sensed Random Measurements
- Learning Attributes Equals Multi-Source Domain Generalization
- Beyond Frontal Faces: Improving Person Recognition Using Multiple Cues
- Don't Just Listen, Use Your Imagination: Leveraging Visual Common Sense for Non-Visual Tasks
- Multi-Cue Zero-Shot Learning with Strong Supervision
- Dense Optical Flow Prediction from a Static Image
- DeepEdge: A Multi-Scale Bifurcated Deep Network for Top-Down Contour Detection
- Class Rectification Hard Mining for Imbalanced Deep Learning
- Part-Stacked CNN for Fine-Grained Visual Categorization
- Parsing Occluded People by Flexible Compositions
- Query-free Clothing Retrieval via Implicit Relevance Feedback