Deep multi-task learning for a geographically-regularized semantic segmentation of aerial images
arXiv:1808.07675 · doi:10.1016/j.isprsjprs.2018.06.007
Abstract
When approaching the semantic segmentation of overhead imagery in the decimeter spatial resolution range, successful strategies usually combine powerful methods to learn the visual appearance of the semantic classes (e.g. convolutional neural networks) with strategies for spatial regularization (e.g. graphical models such as conditional random fields). In this paper, we propose a method to learn evidence in the form of semantic class likelihoods, semantic boundaries across classes and shallow-to-deep visual features, each one modeled by a multi-task convolutional neural network architecture. We combine this bottom-up information with top-down spatial regularization encoded by a conditional random field model optimizing the label space across a hierarchy of segments with constraints related to structural, spatial and data-dependent pairwise relationships between regions. Our results show that such strategy provide better regularization than a series of strong baselines reflecting state-of-the-art technologies. The proposed strategy offers a flexible and principled framework to include several sources of visual and structural information, while allowing for different degrees of spatial regularization accounting for priors about the expected output structures.
References in corpus (6)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Fully Convolutional Networks for Semantic Segmentation
- Dense semantic labeling of sub-decimeter resolution images with convolutional neural networks
- Land cover mapping at very high resolution with rotation equivariant CNNs: towards small yet accurate models
- Fully Convolutional Networks for Dense Semantic Labelling of High-Resolution Aerial Imagery
- High-Resolution Semantic Labeling with Convolutional Neural Networks
Cited by in corpus (8)
- More Diverse Means Better: Multimodal Deep Learning Meets Remote Sensing Imagery Classification
- Understanding urban landuse from the above and ground perspectives: a deep learning, multimodal solution
- Aerial Imagery for Roof Segmentation: A Large-Scale Dataset towards Automatic Mapping of Buildings
- Building Footprint Generation by IntegratingConvolution Neural Network with Feature PairwiseConditional Random Field (FPCRF)
- Robust building footprint extraction from big multi-sensor data using deep competition network
- Boundary Regularized Building Footprint Extraction From Satellite Images Using Deep Neural Network
- Robust object extraction from remote sensing data
- A hierarchical deep learning framework for the consistent classification of land use objects in geospatial databases