Deep Learning for Semantic Part Segmentation with High-Level Guidance
arXiv:1505.02438
Abstract
In this work we address the task of segmenting an object into its parts, or semantic part segmentation. We start by adapting a state-of-the-art semantic segmentation system to this task, and show that a combination of a fully-convolutional Deep CNN system coupled with Dense CRF labelling provides excellent results for a broad range of object categories. Still, this approach remains agnostic to high-level constraints between object parts. We introduce such prior information by means of the Restricted Boltzmann Machine, adapted to our task and train our model in an discriminative fashion, as a hidden CRF, demonstrating that prior information can yield additional improvements. We also investigate the performance of our approach ``in the wild'', without information concerning the objects' bounding boxes, using an object detector to guide a multi-scale segmentation scheme. We evaluate the performance of our approach on the Penn-Fudan and LFW datasets for the tasks of pedestrian parsing and face labelling respectively. We show superior performance with respect to competitive methods that have been extensively engineered on these benchmarks, as well as realistic qualitative results on part segmentation, even for occluded or deformable objects. We also provide quantitative and extensive qualitative results on three classes from the PASCAL Parts dataset. Finally, we show that our multi-scale segmentation scheme can boost accuracy, recovering segmentations for finer parts.
11 pages (including references), 3 figures, 2 tables
References in corpus (9)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Going Deeper with Convolutions
- Fully Convolutional Networks for Semantic Segmentation
- Fully Connected Deep Structured Networks
- Conditional Restricted Boltzmann Machines for Structured Output Prediction
- Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts
- Untangling Local and Global Deformations in Deep Convolutional Networks for Image Classification and Sliding Window Detection
- Hypercolumns for Object Segmentation and Fine-grained Localization
Cited by in corpus (14)
- Convolutional Neural Fabrics
- Illuminating Pedestrians via Simultaneous Detection & Segmentation
- Gated Feedback Refinement Network for Coarse-to-Fine Dense Semantic Image Labeling
- Zoom Better to See Clearer: Human and Object Parsing with Hierarchical Auto-Zoom Net
- Growing Interpretable Part Graphs on ConvNets via Multi-Shot Learning
- Pose-Guided Human Parsing with Deep Learned Features
- Face Parsing via Recurrent Propagation
- Act the Part: Learning Interaction Strategies for Articulated Object Part Discovery
- Face Parsing with RoI Tanh-Warping
- Structured Output Learning with Conditional Generative Flows
- Keypoint Based Weakly Supervised Human Parsing
- Unsupervised Part Discovery from Contrastive Reconstruction
- Learning Discriminators as Energy Networks in Adversarial Learning
- A CNN Cascade for Landmark Guided Semantic Part Segmentation