Weakly Supervised Semantic Segmentation using Web-Crawled Videos
arXiv:1701.00352
Abstract
We propose a novel algorithm for weakly supervised semantic segmentation based on image-level class labels only. In weakly supervised setting, it is commonly observed that trained model overly focuses on discriminative parts rather than the entire object area. Our goal is to overcome this limitation with no additional human intervention by retrieving videos relevant to target class labels from web repository, and generating segmentation labels from the retrieved videos to simulate strong supervision for semantic segmentation. During this process, we take advantage of image classification with discriminative localization technique to reject false alarms in retrieved videos and identify relevant spatio-temporal volumes within retrieved videos. Although the entire procedure does not require any additional supervision, the segmentation annotations obtained from videos are sufficiently strong to learn a model for semantic segmentation. The proposed algorithm substantially outperforms existing methods based on the same level of supervision and is even as competitive as the approaches relying on extra annotations.
CVPR 2017 (Spotlight)
References in corpus (6)
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Learning Deconvolution Network for Semantic Segmentation
- Fully Convolutional Multi-Class Multiple Instance Learning
- Semantic Image Segmentation via Deep Parsing Network
- BoxSup: Exploiting Bounding Boxes to Supervise Convolutional Networks for Semantic Segmentation
- Object Detection, Tracking, and Motion Segmentation for Object-level Video Segmentation
Cited by in corpus (10)
- Revisiting Dilated Convolution: A Simple Approach for Weakly- and Semi- Supervised Semantic Segmentation
- Identity-Guided Human Semantic Parsing for Person Re-Identification
- CEREALS - Cost-Effective REgion-based Active Learning for Semantic Segmentation
- Regularizing Proxies with Multi-Adversarial Training for Unsupervised Domain-Adaptive Semantic Segmentation
- Decoupled Spatial Neural Attention for Weakly Supervised Semantic Segmentation
- Mix-and-Match Tuning for Self-Supervised Semantic Segmentation
- Coarse-to-fine Semantic Segmentation from Image-level Labels
- Block Annotation: Better Image Annotation for Semantic Segmentation with Sub-Image Decomposition
- Concept Mask: Large-Scale Segmentation from Semantic Concepts
- AutoLoc: Weakly-supervised Temporal Action Localization