Diverse Sampling for Self-Supervised Learning of Semantic Segmentation
arXiv:1612.01991
Abstract
We propose an approach for learning category-level semantic segmentation purely from image-level classification tags indicating presence of categories. It exploits localization cues that emerge from training classification-tasked convolutional networks, to drive a "self-supervision" process that automatically labels a sparse, diverse training set of points likely to belong to classes of interest. Our approach has almost no hyperparameters, is modular, and allows for very fast training of segmentation in less than 3 minutes. It obtains competitive results on the VOC 2012 segmentation benchmark. More, significantly the modularity and fast training of our framework allows new classes to efficiently added for inference.
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Object Detectors Emerge in Deep Scene CNNs
- Fully Convolutional Multi-Class Multiple Instance Learning
- BoxSup: Exploiting Bounding Boxes to Supervise Convolutional Networks for Semantic Segmentation
- Feedforward semantic segmentation with zoom-out features
- From Image-level to Pixel-level Labeling with Convolutional Networks