Gated CRF Loss for Weakly Supervised Semantic Image Segmentation
arXiv:1906.04651
Abstract
State-of-the-art approaches for semantic segmentation rely on deep convolutional neural networks trained on fully annotated datasets, that have been shown to be notoriously expensive to collect, both in terms of time and money. To remedy this situation, weakly supervised methods leverage other forms of supervision that require substantially less annotation effort, but they typically present an inability to predict precise object boundaries due to approximate nature of the supervisory signals in those regions. While great progress has been made in improving the performance, many of these weakly supervised methods are highly tailored to their own specific settings. This raises challenges in reusing algorithms and making steady progress. In this paper, we intentionally avoid such practices when tackling weakly supervised semantic segmentation. In particular, we train standard neural networks with partial cross-entropy loss function for the labeled pixels and our proposed Gated CRF loss for the unlabeled pixels. The Gated CRF loss is designed to deliver several important assets: 1) it enables flexibility in the kernel construction to mask out influence from undesired pixel positions; 2) it offloads learning contextual relations to CNN and concentrates on semantic boundaries; 3) it does not rely on high-dimensional filtering and thus has a simple implementation. Throughout the paper we present the advantages of the loss function, analyze several aspects of weakly supervised training, and show that our `purist' approach achieves state-of-the-art performance for both click-based and scribble-based annotations.
A portion of the reported numbers are incorrect, along with a few statements
References in corpus (7)
- Rethinking Atrous Convolution for Semantic Image Segmentation
- Efficient Inference in Fully Connected CRFs with Gaussian Edge Potentials
- Fully Convolutional Multi-Class Multiple Instance Learning
- ICNet for Real-Time Semantic Segmentation on High-Resolution Images
- Rethinking ImageNet Pre-training
- Normalized Cut Loss for Weakly-supervised CNN Segmentation
- On Regularized Losses for Weakly-supervised CNN Segmentation
Cited by in corpus (9)
- PyMIC: A deep learning toolkit for annotation-efficient medical image segmentation
- A Visual Representation-guided Framework with Global Affinity for Weakly Supervised Salient Object Detection
- Synthesize Boundaries: A Boundary-aware Self-consistent Framework for Weakly Supervised Salient Object Detection
- Structure-Consistent Weakly Supervised Salient Object Detection with Local Saliency Coherence
- Weakly-Supervised Salient Object Detection via Scribble Annotations
- Dynamic Feature Regularized Loss for Weakly Supervised Semantic Segmentation
- Scribble-based Weakly Supervised Deep Learning for Road Surface Extraction from Remote Sensing Images
- MambaEviScrib: Mamba and Evidence-Guided Consistency Enhance CNN Robustness for Scribble-Based Weakly Supervised Ultrasound Image Segmentation
- WeClick: Weakly-Supervised Video Semantic Segmentation with Click Annotations