Deep Human Parsing with Active Template Regression
arXiv:1503.02391 · doi:10.1109/TPAMI.2015.2408360
Abstract
In this work, the human parsing task, namely decomposing a human image into semantic fashion/body regions, is formulated as an Active Template Regression (ATR) problem, where the normalized mask of each fashion/body item is expressed as the linear combination of the learned mask templates, and then morphed to a more precise mask with the active shape parameters, including position, scale and visibility of each semantic region. The mask template coefficients and the active shape parameters together can generate the human parsing results, and are thus called the structure outputs for human parsing. The deep Convolutional Neural Network (CNN) is utilized to build the end-to-end relation between the input human image and the structure outputs for human parsing. More specifically, the structure outputs are predicted by two separate networks. The first CNN network is with max-pooling, and designed to predict the template coefficients for each label mask, while the second CNN network is without max-pooling to preserve sensitivity to label mask position and accurately predict the active shape parameters. For a new image, the structure outputs of the two networks are fused to generate the probability of each label for each pixel, and super-pixel smoothing is finally used to refine the human parsing result. Comprehensive evaluations on a large dataset well demonstrate the significant superiority of the ATR framework over other state-of-the-arts for human parsing. In particular, the F1-score reaches by our ATR framework, significantly higher than based on the state-of-the-art algorithm.
This manuscript is the accepted version for IEEE Transactions on Pattern Analysis and Machine Intelligence (TPAMI) 2015
References in corpus (2)
Cited by in corpus (51)
- Semantic-guided Pixel Sampling for Cloth-Changing Person Re-identification
- Makeup like a superstar: Deep Localized Makeup Transfer Network
- Look into Person: Self-supervised Structure-sensitive Learning and A New Benchmark for Human Parsing
- Fashion Meets Computer Vision: A Survey
- Looking at Outfit to Parse Clothing
- A Comprehensive Review of Modern Object Segmentation Approaches
- Hough-CNN: Deep Learning for Segmentation of Deep Brain Regions in MRI and Ultrasound
- Semantic Object Parsing with Local-Global Long Short-Term Memory
- Instance-level Human Parsing via Part Grouping Network
- ModaNet: A Large-Scale Street Fashion Dataset with Polygon Annotations
- Deep Learning Technique for Human Parsing: A Survey and Outlook
- TextureGAN: Controlling Deep Image Synthesis with Texture Patches
- Matching-CNN Meets KNN: Quasi-Parametric Human Parsing
- Holistic, Instance-Level Human Parsing
- ViTAA: Visual-Textual Attributes Alignment in Person Search by Natural Language
- Textured Neural Avatars
- Cross-domain Human Parsing via Adversarial Feature and Label Adaptation
- Learning Deep Representations for Semantic Image Parsing: a Comprehensive Overview
- Deep Co-Space: Sample Mining Across Feature Transformation for Semi-Supervised Learning
- Hierarchical Human Parsing with Typed Part-Relation Reasoning
- fAshIon after fashion: A Report of AI in Fashion
- High-Resolution Image Inpainting with Iterative Confidence Feedback and Guided Upsampling
- Semantic Object Parsing with Graph LSTM
- Surveillance Video Parsing with Single Frame Supervision
- Learning Compositional Neural Information Fusion for Human Parsing
- CRAFT: Complementary Recommendations Using Adversarial Feature Transformer
- Apparel-invariant Feature Learning for Apparel-changed Person Re-identification
- Multi-Task Curriculum Transfer Deep Learning of Clothing Attributes
- Human Pose Estimation from Depth Images via Inference Embedded Multi-task Learning
- Integrated Inference and Learning of Neural Factors in Structural Support Vector Machines
- 3D Magic Mirror: Clothing Reconstruction from a Single Image via a Causal Perspective
- Spatial-Aware Non-Local Attention for Fashion Landmark Detection
- Arbitrary Virtual Try-On Network: Characteristics Preservation and Trade-off between Body and Clothing
- Controllable Image Synthesis via SegVAE
- Adaptive Temporal Encoding Network for Video Instance-level Human Parsing
- Unselfie: Translating Selfies to Neutral-pose Portraits in the Wild
- Attention-based fusion of semantic boundary and non-boundary information to improve semantic segmentation
- Graphonomy: Universal Image Parsing via Graph Reasoning and Transfer
- End-to-end One-shot Human Parsing
- Body Segmentation Using Multi-task Learning
- Fast Soft Color Segmentation
- Human-centric Relation Segmentation: Dataset and Solution
- Deep Image Synthesis from Intuitive User Input: A Review and Perspectives
- Self-Learning with Rectification Strategy for Human Parsing
- Progressive refinement: a method of coarse-to-fine image parsing using stacked network
- Bandwidth limited object recognition in high resolution imagery
- Multi-Attribute Enhancement Network for Person Search
- Video Deblurring via Semantic Segmentation and Pixel-Wise Non-Linear Kernel
- Learning Intrinsic Images for Clothing
- Digital Makeup from Internet Images
- Tracking by Joint Local and Global Search: A Target-aware Attention based Approach