STC: A Simple to Complex Framework for Weakly-supervised Semantic Segmentation
arXiv:1509.03150 · doi:10.1109/TPAMI.2016.2636150
Abstract
Recently, significant improvement has been made on semantic object segmentation due to the development of deep convolutional neural networks (DCNNs). Training such a DCNN usually relies on a large number of images with pixel-level segmentation masks, and annotating these images is very costly in terms of both finance and human effort. In this paper, we propose a simple to complex (STC) framework in which only image-level annotations are utilized to learn DCNNs for semantic segmentation. Specifically, we first train an initial segmentation network called Initial-DCNN with the saliency maps of simple images (i.e., those with a single category of major object(s) and clean background). These saliency maps can be automatically obtained by existing bottom-up salient object detection techniques, where no supervision information is needed. Then, a better network called Enhanced-DCNN is learned with supervision from the predicted segmentation masks of simple images based on the Initial-DCNN as well as the image-level annotations. Finally, more pixel-level segmentation masks of complex images (two or more categories of objects with cluttered background), which are inferred by using Enhanced-DCNN and image-level annotations, are utilized as the supervision information to learn the Powerful-DCNN for semantic segmentation. Our method utilizes K simple images from Flickr.com and 10K complex images from PASCAL VOC for step-wisely boosting the segmentation network. Extensive experimental results on PASCAL VOC 2012 segmentation benchmark well demonstrate the superiority of the proposed STC framework compared with other state-of-the-arts.
To Appear in IEEE Transactions on Pattern Analysis and Machine Intelligence
References in corpus (11)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Semantic Image Segmentation with Deep Convolutional Nets and Fully Connected CRFs
- Fully Convolutional Networks for Semantic Segmentation
- Conditional Random Fields as Recurrent Neural Networks
- Salient Object Detection: A Benchmark
- Convolutional Feature Masking for Joint Object and Stuff Segmentation
- Weakly- and Semi-Supervised Learning of a DCNN for Semantic Image Segmentation
- Fast R-CNN
- Fully Convolutional Multi-Class Multiple Instance Learning
- Computational Baby Learning
Cited by in corpus (97)
- The Cityscapes Dataset for Semantic Urban Scene Understanding
- Richer Convolutional Features for Edge Detection
- Deeply supervised salient object detection with short connections
- Constrained-CNN losses for weakly supervised segmentation
- Weakly Supervised Adversarial Domain Adaptation for Semantic Segmentation in Urban Scenes
- CAVER: Cross-Modal View-Mixed Transformer for Bi-Modal Salient Object Detection
- Few-Example Object Detection with Model Communication
- Affinity Attention Graph Neural Network for Weakly Supervised Semantic Segmentation
- Edge Preserving and Multi-Scale Contextual Neural Network for Salient Object Detection
- Self Paced Deep Learning for Weakly Supervised Object Detection
- Light Field Salient Object Detection: A Review and Benchmark
- Weakly-Supervised Semantic Segmentation by Iterative Affinity Learning
- SG-One: Similarity Guidance Network for One-Shot Semantic Segmentation
- Large-scale Unsupervised Semantic Segmentation
- Cross-layer Feature Pyramid Network for Salient Object Detection
- Adversarial Complementary Learning for Weakly Supervised Object Localization
- Tell Me Where to Look: Guided Attention Inference Network
- Saliency Guided Inter- and Intra-Class Relation Constraints for Weakly Supervised Semantic Segmentation
- Devil in the Details: Towards Accurate Single and Multiple Human Parsing
- Consistency-Regularized Region-Growing Network for Semantic Segmentation of Urban Scenes with Point-Level Annotations
- Reverse Attention for Salient Object Detection
- Self-supervised Equivariant Attention Mechanism for Weakly Supervised Semantic Segmentation
- Gaussian Dynamic Convolution for Efficient Single-Image Segmentation
- Digital Twin-based Anomaly Detection with Curriculum Learning in Cyber-physical Systems
- Object Region Mining with Adversarial Erasing: A Simple Classification to Semantic Segmentation Approach
- Multi-Granularity Denoising and Bidirectional Alignment for Weakly Supervised Semantic Segmentation
- Revisiting Dilated Convolution: A Simple Approach for Weakly- and Semi- Supervised Semantic Segmentation
- Salient Objects in Clutter: Bringing Salient Object Detection to the Foreground
- Learning to Exploit the Prior Network Knowledge for Weakly-Supervised Semantic Segmentation
- Incorporating Network Built-in Priors in Weakly-supervised Semantic Segmentation
- Two-Phase Learning for Weakly Supervised Object Localization
- Weakly Supervised Instance Segmentation using Class Peak Response
- Weakly-Supervised Semantic Segmentation by Iteratively Mining Common Object Features
- Structure Label Prediction Using Similarity-Based Retrieval and Weakly-Supervised Label Mapping
- Exploiting saliency for object segmentation from image level labels
- Simple Does It: Weakly Supervised Instance and Semantic Segmentation
- Constrained domain adaptation for Image segmentation
- Built-in Foreground/Background Prior for Weakly-Supervised Semantic Segmentation
- Unsupervised Domain Adaptation in Semantic Segmentation: a Review
- Multi-task deep learning for image segmentation using recursive approximation tasks
- Self-supervised Scale Equivariant Network for Weakly Supervised Semantic Segmentation
- Backtracking Spatial Pyramid Pooling (SPP)-based Image Classifier for Weakly Supervised Top-down Salient Object Detection
- CIAN: Cross-Image Affinity Net for Weakly Supervised Semantic Segmentation
- A Single Stream Network for Robust and Real-time RGB-D Salient Object Detection
- Effective Use of Synthetic Data for Urban Scene Semantic Segmentation
- Group-Wise Semantic Mining for Weakly Supervised Semantic Segmentation
- Regularizing Proxies with Multi-Adversarial Training for Unsupervised Domain-Adaptive Semantic Segmentation
- PseudoAugment: Learning to Use Unlabeled Data for Data Augmentation in Point Clouds
- Non-Salient Region Object Mining for Weakly Supervised Semantic Segmentation
- Hierarchical Dynamic Filtering Network for RGB-D Salient Object Detection
- Cross-domain Human Parsing via Adversarial Feature and Label Adaptation
- Deep-Energy: Unsupervised Training of Deep Neural Networks
- Deep Reasoning with Multi-Scale Context for Salient Object Detection
- Inter-Image Communication for Weakly Supervised Localization
- Naive-Student: Leveraging Semi-Supervised Learning in Video Sequences for Urban Scene Segmentation
- Saliency Guided Self-attention Network for Weakly and Semi-supervised Semantic Segmentation
- Decoupled Spatial Neural Attention for Weakly Supervised Semantic Segmentation
- Segregated Temporal Assembly Recurrent Networks for Weakly Supervised Multiple Action Detection
- Complementary Patch for Weakly Supervised Semantic Segmentation
- Self-produced Guidance for Weakly-supervised Object Localization
- WebSeg: Learning Semantic Segmentation from Web Searches
- Progressively Guided Alternate Refinement Network for RGB-D Salient Object Detection
- Learning structure-aware semantic segmentation with image-level supervision
- Importance Sampling CAMs for Weakly-Supervised Segmentation
- Adaptive Early-Learning Correction for Segmentation from Noisy Annotations
- Exploiting Web Images for Weakly Supervised Object Detection
- Adaptive Affinity Loss and Erroneous Pseudo-Label Refinement for Weakly Supervised Semantic Segmentation
- Weakly Supervised Faster-RCNN+FPN to classify animals in camera trap images
- Confidence-Guided Learning Process for Continuous Classification of Time Series
- Progressive Self-Guided Loss for Salient Object Detection
- Contrast-Oriented Deep Neural Networks for Salient Object Detection
- S4Net: Single Stage Salient-Instance Segmentation
- CLASS: Cross-Level Attention and Supervision for Salient Objects Detection
- Cross-Image Region Mining with Region Prototypical Network for Weakly Supervised Segmentation
- Discovering Latent Classes for Semi-Supervised Semantic Segmentation
- Weakly Supervised Body Part Segmentation with Pose based Part Priors
- Saliency guided deep network for weakly-supervised image segmentation
- Weakly- and Semi-Supervised Panoptic Segmentation
- Diverse Sampling for Self-Supervised Learning of Semantic Segmentation
- W-TALC: Weakly-supervised Temporal Activity Localization and Classification
- Generating Self-Guided Dense Annotations for Weakly Supervised Semantic Segmentation
- Weakly-supervised land classification for coastal zone based on deep convolutional neural networks by incorporating dual-polarimetric characteristics into training dataset
- Gradient-Induced Co-Saliency Detection
- Cognition Transferring and Decoupling for Text-supervised Egocentric Semantic Segmentation
- Automatic Image Labelling at Pixel Level
- Learning Pixel-wise Labeling from the Internet without Human Interaction
- Learning to segment with image-level supervision
- Realizing Pixel-Level Semantic Learning in Complex Driving Scenes based on Only One Annotated Pixel per Class
- ROSA: Robust Salient Object Detection against Adversarial Attacks
- RealCustom++: Representing Images as Real Textual Word for Real-Time Customization
- Closed-Loop Adaptation for Weakly-Supervised Semantic Segmentation
- ACFNet: Adaptively-Cooperative Fusion Network for RGB-D Salient Object Detection
- TS2C: Tight Box Mining with Surrounding Segmentation Context for Weakly Supervised Object Detection
- DASNet: Reducing Pixel-level Annotations for Instance and Semantic Segmentation
- Maximize the Exploration of Congeneric Semantics for Weakly Supervised Semantic Segmentation
- Learning Rich Representations For Structured Visual Prediction Tasks
- Seed, Expand and Constrain: Three Principles for Weakly-Supervised Image Segmentation