Instance-aware Semantic Segmentation via Multi-task Network Cascades
arXiv:1512.04412
Abstract
Semantic segmentation research has recently witnessed rapid progress, but many leading methods are unable to identify object instances. In this paper, we present Multi-task Network Cascades for instance-aware semantic segmentation. Our model consists of three networks, respectively differentiating instances, estimating masks, and categorizing objects. These networks form a cascaded structure, and are designed to share their convolutional features. We develop an algorithm for the nontrivial end-to-end training of this causal, cascaded structure. Our solution is a clean, single-step training framework and can be generalized to cascades that have more stages. We demonstrate state-of-the-art instance-aware semantic segmentation accuracy on PASCAL VOC. Meanwhile, our method takes only 360ms testing an image using VGG-16, which is two orders of magnitude faster than previous systems for this challenging problem. As a by product, our method also achieves compelling object detection results which surpass the competitive Fast/Faster R-CNN systems. The method described in this paper is the foundation of our submissions to the MS COCO 2015 segmentation competition, where we won the 1st place.
Tech report. 1st-place winner of MS COCO 2015 segmentation competition
References in corpus (3)
Cited by in corpus (49)
- FCNs in the Wild: Pixel-level Adversarial and Constraint-based Adaptation
- Semantic Instance Segmentation with a Discriminative Loss Function
- Systematic evaluation of CNN advances on the ImageNet
- What makes ImageNet good for transfer learning?
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- Semantic Instance Segmentation via Deep Metric Learning
- Speed/accuracy trade-offs for modern convolutional object detectors
- Adversarial Networks for Spatial Context-Aware Spectral Image Reconstruction from RGB
- VoxResNet: Deep Voxelwise Residual Networks for Volumetric Brain Segmentation
- ME R-CNN: Multi-Expert R-CNN for Object Detection
- Illuminating Pedestrians via Simultaneous Detection & Segmentation
- A Survey of Semantic Segmentation
- Learning deep structured active contours end-to-end
- A deep architecture for unified aesthetic prediction
- Weakly Supervised Instance Segmentation using Class Peak Response
- Deep Watershed Transform for Instance Segmentation
- Yum-me: A Personalized Nutrient-based Meal Recommender System
- Efficient Two-Stream Motion and Appearance 3D CNNs for Video Classification
- MegDet: A Large Mini-Batch Object Detector
- Distinguishing mirror from glass: A 'big data' approach to material perception
- Instance-level Human Parsing via Part Grouping Network
- InstanceCut: from Edges to Instances with MultiCut
- Learning to Segment Instances in Videos with Spatial Propagation Network
- LabelBank: Revisiting Global Perspectives for Semantic Segmentation
- Deep Cuboid Detection: Beyond 2D Bounding Boxes
- Deep-CEE I: Fishing for Galaxy Clusters with Deep Neural Nets
- Learning to Segment Every Thing
- Deconvolutional Feature Stacking for Weakly-Supervised Semantic Segmentation
- Amodal Instance Segmentation
- Estimated Depth Map Helps Image Classification
- Boundary-aware Instance Segmentation
- Spatial Feature Calibration and Temporal Fusion for Effective One-stage Video Instance Segmentation
- SketchParse : Towards Rich Descriptions for Poorly Drawn Sketches using Multi-Task Hierarchical Deep Networks
- Fine-grained Recognition in the Wild: A Multi-Task Domain Adaptation Approach
- Instance Shadow Detection
- Straight to Shapes: Real-time Detection of Encoded Shapes
- Differentiable Multi-Granularity Human Representation Learning for Instance-Aware Human Semantic Parsing
- Pseudo Mask Augmented Object Detection
- SeGAN: Segmenting and Generating the Invisible
- Dense 3D Regression for Hand Pose Estimation
- Center-Focusing Multi-task CNN with Injected Features for Classification of Glioma Nuclear Images
- Fast and Accurate Online Video Object Segmentation via Tracking Parts
- CASNet: Common Attribute Support Network for image instance and panoptic segmentation
- Instance-Level Salient Object Segmentation
- Learning and Memorizing Representative Prototypes for 3D Point Cloud Semantic and Instance Segmentation
- StuffNet: Using 'Stuff' to Improve Object Detection
- Object Boundary Guided Semantic Segmentation
- VideoClick: Video Object Segmentation with a Single Click
- A Distraction Score for Watermarks