Segmentation Transformer: Object-Contextual Representations for Semantic Segmentation
arXiv:1909.11065
Abstract
In this paper, we address the semantic segmentation problem with a focus on the context aggregation strategy. Our motivation is that the label of a pixel is the category of the object that the pixel belongs to. We present a simple yet effective approach, object-contextual representations, characterizing a pixel by exploiting the representation of the corresponding object class. First, we learn object regions under the supervision of ground-truth segmentation. Second, we compute the object region representation by aggregating the representations of the pixels lying in the object region. Last, % the representation similarity we compute the relation between each pixel and each object region and augment the representation of each pixel with the object-contextual representation which is a weighted aggregation of all the object region representations according to their relations with the pixel. We empirically demonstrate that the proposed approach achieves competitive performance on various challenging semantic segmentation benchmarks: Cityscapes, ADE20K, LIP, PASCAL-Context, and COCO-Stuff. Cityscapes, ADE20K, LIP, PASCAL-Context, and COCO-Stuff. Our submission "HRNet + OCR + SegFix" achieves 1-st place on the Cityscapes leaderboard by the time of submission. Code is available at: https://git.io/openseg and https://git.io/HRNet.OCR. We rephrase the object-contextual representation scheme using the Transformer encoder-decoder framework. The details are presented in~Section3.3.
We rephrase the object-contextual representation scheme using the Transformer encoder-decoder framework. ECCV 2020 Spotlight. Project Page: https://github.com/openseg-group/openseg.pytorch https://github.com/HRNet/HRNet-Semantic-Segmentation/tree/HRNet-OCR
References in corpus (7)
- Rethinking Atrous Convolution for Semantic Image Segmentation
- High-Resolution Representations for Labeling Pixels and Regions
- Hierarchical Multi-Scale Attention for Semantic Segmentation
- Interlaced Sparse Self-Attention for Semantic Segmentation
- Label Refinement Network for Coarse-to-Fine Semantic Segmentation
- Global Aggregation then Local Distribution in Fully Convolutional Networks
- ShapeMask: Learning to Segment Novel Objects by Refining Shape Priors
Cited by in corpus (24)
- Attention Mechanisms in Computer Vision: A Survey
- Radar-Camera Fusion for Object Detection and Semantic Segmentation in Autonomous Driving: A Comprehensive Review
- Building extraction with vision transformer
- HED-UNet: Combined Segmentation and Edge Detection for Monitoring the Antarctic Coastline
- RGB-T Semantic Segmentation with Location, Activation, and Sharpening
- Looking Outside the Window: Wide-Context Transformer for the Semantic Segmentation of High-Resolution Remote Sensing Images
- Large-scale Unsupervised Semantic Segmentation
- Uncertainty-aware Contrastive Distillation for Incremental Semantic Segmentation
- AIParsing: Anchor-free Instance-level Human Parsing
- On the Real-World Adversarial Robustness of Real-Time Semantic Segmentation Models for Autonomous Driving
- Analyzing Green View Index and Green View Index best path using Google Street View and deep learning
- A Comprehensive Review of Modern Object Segmentation Approaches
- Attention guided global enhancement and local refinement network for semantic segmentation
- SBSS: Stacking-Based Semantic Segmentation Framework for Very High Resolution Remote Sensing Image
- Hidden Path Selection Network for Semantic Segmentation of Remote Sensing Images
- Low-Resolution Self-Attention for Semantic Segmentation
- SVCNet: Scribble-based Video Colorization Network with Temporal Aggregation
- Graph-Segmenter: Graph Transformer with Boundary-aware Attention for Semantic Segmentation
- HRNET: AI on Edge for mask detection and social distancing
- Deep Common Feature Mining for Efficient Video Semantic Segmentation
- Which cycling environment appears safer? Learning cycling safety perceptions from pairwise image comparisons
- Global and Local Features through Gaussian Mixture Models on Image Semantic Segmentation
- Continual Learning: Forget-free Winning Subnetworks for Video Representations
- SegAssess: Panoramic quality mapping for robust and transferable unsupervised segmentation assessment