Zoom Better to See Clearer: Human and Object Parsing with Hierarchical Auto-Zoom Net
arXiv:1511.06881
Abstract
Parsing articulated objects, e.g. humans and animals, into semantic parts (e.g. body, head and arms, etc.) from natural images is a challenging and fundamental problem for computer vision. A big difficulty is the large variability of scale and location for objects and their corresponding parts. Even limited mistakes in estimating scale and location will degrade the parsing output and cause errors in boundary details. To tackle these difficulties, we propose a "Hierarchical Auto-Zoom Net" (HAZN) for object part parsing which adapts to the local scales of objects and parts. HAZN is a sequence of two "Auto-Zoom Net" (AZNs), each employing fully convolutional networks that perform two tasks: (1) predict the locations and scales of object instances (the first AZN) or their parts (the second AZN); (2) estimate the part scores for predicted object instance or part regions. Our model can adaptively "zoom" (resize) predicted image regions into their proper scales to refine the parsing. We conduct extensive experiments over the PASCAL part datasets on humans, horses, and cows. For humans, our approach significantly outperforms the state-of-the-arts by 5% mIOU and is especially better at segmenting small instances and small parts. We obtain similar improvements for parsing cows and horses over alternative methods. In summary, our strategy of first zooming into objects and then zooming into parts is very effective. It also enables us to process different regions of the image at different scales adaptively so that, for example, we do not need to waste computational resources scaling the entire image.
A shortened version has been submitted to ECCV 2016
References in corpus (14)
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Learning Deconvolution Network for Semantic Segmentation
- DenseBox: Unifying Landmark Localization with End to End Object Detection
- Semantic Image Segmentation via Deep Parsing Network
- BoxSup: Exploiting Bounding Boxes to Supervise Convolutional Networks for Semantic Segmentation
- Proposal-free Network for Instance-level Object Segmentation
- Detect What You Can: Detecting and Representing Objects using Holistic Models and Body Parts
- segDeepM: Exploiting Segmentation and Context in Deep Neural Networks for Object Detection
- Deep Learning for Semantic Part Segmentation with High-Level Guidance
- Joint Object and Part Segmentation using Deep Learned Potentials
- Matching-CNN Meets KNN: Quasi-Parametric Human Parsing
- Amodal Completion and Size Constancy in Natural Scenes
- Pose-Guided Human Parsing with Deep Learned Features
Cited by in corpus (6)
- DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
- RefineNet: Multi-Path Refinement Networks for High-Resolution Semantic Segmentation
- Cross-domain Human Parsing via Adversarial Feature and Label Adaptation
- Surveillance Video Parsing with Single Frame Supervision
- Semantic Object Parsing with Graph LSTM
- DeepSkeleton: Skeleton Map for 3D Human Pose Regression