Contrastive Multiview Coding with Electro-optics for SAR Semantic Segmentation
arXiv:2109.00120 · doi:10.1109/LGRS.2021.3109345
Abstract
In the training of deep learning models, how the model parameters are initialized greatly affects the model performance, sample efficiency, and convergence speed. Representation learning for model initialization has recently been actively studied in the remote sensing field. In particular, the appearance characteristics of the imagery obtained using the a synthetic aperture radar (SAR) sensor are quite different from those of general electro-optical (EO) images, and thus representation learning is even more important in remote sensing domain. Motivated from contrastive multiview coding, we propose multi-modal representation learning for SAR semantic segmentation. Unlike previous studies, our method jointly uses EO imagery, SAR imagery, and a label mask. Several experiments show that our approach is superior to the existing methods in model performance, sample efficiency, and convergence speed.
To be appeared in IEEE GRSL. DOI to be updated
References in corpus (4)
- Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
- Rethinking Atrous Convolution for Semantic Image Segmentation
- SpaceNet 6: Multi-Sensor All Weather Mapping Dataset
- Revisiting Classical Bagging with Modern Transfer Learning for On-the-fly Disaster Damage Detector