Fully Convolutional Adaptation Networks for Semantic Segmentation
arXiv:1804.08286
Abstract
The recent advances in deep neural networks have convincingly demonstrated high capability in learning vision models on large datasets. Nevertheless, collecting expert labeled datasets especially with pixel-level annotations is an extremely expensive process. An appealing alternative is to render synthetic data (e.g., computer games) and generate ground truth automatically. However, simply applying the models learnt on synthetic images may lead to high generalization error on real images due to domain shift. In this paper, we facilitate this issue from the perspectives of both visual appearance-level and representation-level domain adaptation. The former adapts source-domain images to appear as if drawn from the "style" in the target domain and the latter attempts to learn domain-invariant representations. Specifically, we present Fully Convolutional Adaptation Networks (FCAN), a novel deep architecture for semantic segmentation which combines Appearance Adaptation Networks (AAN) and Representation Adaptation Networks (RAN). AAN learns a transformation from one domain to the other in the pixel space and RAN is optimized in an adversarial learning manner to maximally fool the domain discriminator with the learnt source and target representations. Extensive experiments are conducted on the transfer from GTA5 (game videos) to Cityscapes (urban street scenes) on semantic segmentation and our proposal achieves superior results when comparing to state-of-the-art unsupervised adaptation techniques. More remarkably, we obtain a new record: mIoU of 47.5% on BDDS (drive-cam videos) in an unsupervised setting.
CVPR 2018, Rank 1 in Segmentation Track of Visual Domain Adaptation Challenge 2017
References in corpus (11)
- Caffe: Convolutional Architecture for Fast Feature Embedding
- Learning Transferable Features with Deep Adaptation Networks
- Deep Domain Confusion: Maximizing for Domain Invariance
- ParseNet: Looking Wider to See Better
- The Cityscapes Dataset for Semantic Urban Scene Understanding
- FCNs in the Wild: Pixel-level Adversarial and Constraint-based Adaptation
- Adversarial Discriminative Domain Adaptation
- Texture Synthesis Using Convolutional Neural Networks
- Fully Convolutional Multi-Class Multiple Instance Learning
- Playing for Data: Ground Truth from Computer Games
- BoxSup: Exploiting Bounding Boxes to Supervise Convolutional Networks for Semantic Segmentation
Cited by in corpus (16)
- A Review on Generative Adversarial Networks: Algorithms, Theory, and Applications
- Adversarial Style Mining for One-Shot Unsupervised Domain Adaptation
- Transferrable Prototypical Networks for Unsupervised Domain Adaptation
- Exploring Object Relation in Mean Teacher for Cross-Domain Detection
- Rectifying Pseudo Label Learning via Uncertainty Estimation for Domain Adaptive Semantic Segmentation
- A Novel Upsampling and Context Convolution for Image Semantic Segmentation
- Regularizing Proxies with Multi-Adversarial Training for Unsupervised Domain-Adaptive Semantic Segmentation
- ACE: Adapting to Changing Environments for Semantic Segmentation
- Customizable Architecture Search for Semantic Segmentation
- Label-Driven Reconstruction for Domain Adaptation in Semantic Segmentation
- Multi-Source Domain Adaptation and Semi-Supervised Domain Adaptation with Focus on Visual Domain Adaptation Challenge 2019
- Alleviating Semantic-level Shift: A Semi-supervised Domain Adaptation Method for Semantic Segmentation
- Transferring and Regularizing Prediction for Semantic Segmentation
- Exploring Category-Agnostic Clusters for Open-Set Domain Adaptation
- Bi-Directional Generation for Unsupervised Domain Adaptation
- Softer Pruning, Incremental Regularization