Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers
arXiv:2012.15840
Abstract
Most recent semantic segmentation methods adopt a fully-convolutional network (FCN) with an encoder-decoder architecture. The encoder progressively reduces the spatial resolution and learns more abstract/semantic visual concepts with larger receptive fields. Since context modeling is critical for segmentation, the latest efforts have been focused on increasing the receptive field, through either dilated/atrous convolutions or inserting attention modules. However, the encoder-decoder based FCN architecture remains unchanged. In this paper, we aim to provide an alternative perspective by treating semantic segmentation as a sequence-to-sequence prediction task. Specifically, we deploy a pure transformer (ie, without convolution and resolution reduction) to encode an image as a sequence of patches. With the global context modeled in every layer of the transformer, this encoder can be combined with a simple decoder to provide a powerful segmentation model, termed SEgmentation TRansformer (SETR). Extensive experiments show that SETR achieves new state of the art on ADE20K (50.28% mIoU), Pascal Context (55.83% mIoU) and competitive results on Cityscapes. Particularly, we achieve the first position in the highly competitive ADE20K test server leaderboard on the day of submission.
CVPR 2021. Project page at https://fudan-zvg.github.io/SETR/
References in corpus (2)
Cited by in corpus (12)
- BEiT: BERT Pre-Training of Image Transformers
- Beyond Self-attention: External Attention using Two Linear Layers for Visual Tasks
- Point-BERT: Pre-training 3D Point Cloud Transformers with Masked Point Modeling
- Unifying Global-Local Representations in Salient Object Detection with Transformer
- TransLoc3D : Point Cloud based Large-scale Place Recognition using Adaptive Receptive Fields
- Multi-Scale Feature Aggregation by Cross-Scale Pixel-to-Region Relation Operation for Semantic Segmentation
- Vision Transformer using Low-level Chest X-ray Feature Corpus for COVID-19 Diagnosis and Severity Quantification
- Tsformer: Time series Transformer for tourism demand forecasting
- RAMS-Trans: Recurrent Attention Multi-scale Transformer forFine-grained Image Recognition
- A Large-Scale Benchmark for Food Image Segmentation
- OadTR: Online Action Detection with Transformers
- Synthesizing Photorealistic Images with Deep Generative Learning