BoxSup: Exploiting Bounding Boxes to Supervise Convolutional Networks for Semantic Segmentation
arXiv:1503.01640
Abstract
Recent leading approaches to semantic segmentation rely on deep convolutional networks trained with human-annotated, pixel-level segmentation masks. Such pixel-accurate supervision demands expensive labeling effort and limits the performance of deep networks that usually benefit from more training data. In this paper, we propose a method that achieves competitive accuracy but only requires easily obtained bounding box annotations. The basic idea is to iterate between automatically generating region proposals and training convolutional networks. These two steps gradually recover segmentation masks for improving the networks, and vise versa. Our method, called BoxSup, produces competitive results supervised by boxes only, on par with strong baselines fully supervised by masks under the same setting. By leveraging a large amount of bounding boxes, BoxSup further unleashes the power of deep convolutional networks and yields state-of-the-art results on PASCAL VOC 2012 and PASCAL-CONTEXT.
References in corpus (5)
Cited by in corpus (22)
- Wider or Deeper: Revisiting the ResNet Model for Visual Recognition
- RefineNet: Multi-Path Refinement Networks for High-Resolution Semantic Segmentation
- Semi and Weakly Supervised Semantic Segmentation Using Generative Adversarial Network
- WordSup: Exploiting Word Annotations for Character based Text Detection
- Built-in Foreground/Background Prior for Weakly-Supervised Semantic Segmentation
- LabelBank: Revisiting Global Perspectives for Semantic Segmentation
- Full-Resolution Residual Networks for Semantic Segmentation in Street Scenes
- Scene Parsing with Global Context Embedding
- Multi-level Contextual RNNs with Attention Model for Scene Labeling
- Accurate Weakly Supervised Deep Lesion Segmentation on CT Scans: Self-Paced 3D Mask Generation from RECIST
- Recent Advances in the Applications of Convolutional Neural Networks to Medical Image Contour Detection
- FoveaNet: Perspective-aware Urban Scene Parsing
- Dense Recurrent Neural Networks for Scene Labeling
- Surveillance Video Parsing with Single Frame Supervision
- Mix-and-Match Tuning for Self-Supervised Semantic Segmentation
- Learning Deep Representations for Scene Labeling with Semantic Context Guided Supervision
- Diverse Sampling for Self-Supervised Learning of Semantic Segmentation
- Recalling Holistic Information for Semantic Segmentation
- weedNet: Dense Semantic Weed Classification Using Multispectral Images and MAV for Smart Farming
- Proposal Flow: Semantic Correspondences from Object Proposals
- Region-based semantic segmentation with end-to-end training
- A Holistic Approach for Data-Driven Object Cutout