Learning to Segment Every Thing
arXiv:1711.10370
Abstract
Most methods for object instance segmentation require all training examples to be labeled with segmentation masks. This requirement makes it expensive to annotate new categories and has restricted instance segmentation models to ~100 well-annotated classes. The goal of this paper is to propose a new partially supervised training paradigm, together with a novel weight transfer function, that enables training instance segmentation models on a large set of categories all of which have box annotations, but only a small fraction of which have mask annotations. These contributions allow us to train Mask R-CNN to detect and segment 3000 visual concepts using box annotations from the Visual Genome dataset and mask annotations from the 80 classes in the COCO dataset. We evaluate our approach in a controlled study on the COCO dataset. This work is a first step towards instance segmentation models that have broad comprehension of the visual world.
References in corpus (6)
- A Learned Representation For Artistic Style
- Feature Pyramid Networks for Object Detection
- BoxSup: Exploiting Bounding Boxes to Supervise Convolutional Networks for Semantic Segmentation
- InstanceCut: from Edges to Instances with MultiCut
- Learning Robust Visual-Semantic Embeddings
- Boundary-aware Instance Segmentation
Cited by in corpus (7)
- CANet: Class-Agnostic Segmentation Networks with Iterative Refinement and Attentive Few-Shot Learning
- Unified Perceptual Parsing for Scene Understanding
- Evaluating Large-Vocabulary Object Detectors: The Devil is in the Details
- Deep Learning Based Instance Segmentation in 3D Biomedical Images Using Weak Annotation
- Weakly- and Semi-Supervised Panoptic Segmentation
- Unseen Object Segmentation in Videos via Transferable Representations
- Meta R-CNN : Towards General Solver for Instance-level Few-shot Learning