Understanding urban landuse from the above and ground perspectives: a deep learning, multimodal solution
arXiv:1905.01752 · doi:10.1016/j.rse.2019.04.014
Abstract
Landuse characterization is important for urban planning. It is traditionally performed with field surveys or manual photo interpretation, two practices that are time-consuming and labor-intensive. Therefore, we aim to automate landuse mapping at the urban-object level with a deep learning approach based on data from multiple sources (or modalities). We consider two image modalities: overhead imagery from Google Maps and ensembles of ground-based pictures (side-views) per urban-object from Google Street View (GSV). These modalities bring complementary visual information pertaining to the urban-objects. We propose an end-to-end trainable model, which uses OpenStreetMap annotations as labels. The model can accommodate a variable number of GSV pictures for the ground-based branch and can also function in the absence of ground pictures at prediction time. We test the effectiveness of our model over the area of Île-de-France, France, and test its generalization abilities on a set of urban-objects from the city of Nantes, France. Our proposed multimodal Convolutional Neural Network achieves considerably higher accuracies than methods that use a single image modality, making it suitable for automatic landuse map updates. Additionally, our approach could be easily scaled to multiple cities, because it is based on data sources available for many cities worldwide.
References in corpus (7)
- Very Deep Convolutional Networks for Large-Scale Image Recognition
- Deep learning in remote sensing: a review
- Dense semantic labeling of sub-decimeter resolution images with convolutional neural networks
- Semisupervised Manifold Alignment of Multimodal Remote Sensing Images
- Multiclass feature learning for hyperspectral image classification: sparse and hierarchical solutions
- Towards seamless multi-view scene analysis from satellite to street-level
- Multi-temporal and multi-source remote sensing image classification by nonlinear relative normalization
Cited by in corpus (17)
- More Diverse Means Better: Multimodal Deep Learning Meets Remote Sensing Imagery Classification
- X-ModalNet: A Semi-Supervised Deep Cross-Modal Network for Classification of Remote Sensing Data
- OpenStreetMap: Challenges and Opportunities in Machine Learning and Remote Sensing
- Enabling Country-Scale Land Cover Mapping with Meter-Resolution Satellite Imagery
- Towards a Collective Agenda on AI for Earth Science Data Analysis
- Dual-Branch Subpixel-Guided Network for Hyperspectral Image Classification
- Common Practices and Taxonomy in Deep Multi-view Fusion for Remote Sensing Applications
- Multisource Collaborative Domain Generalization for Cross-Scene Remote Sensing Image Classification
- Deploying machine learning to assist digital humanitarians: making image annotation in OpenStreetMap more efficient
- Cross-Modal Pre-Aligned Method with Global and Local Information for Remote-Sensing Image and Text Retrieval
- Impact Assessment of Missing Data in Model Predictions for Earth Observation Applications
- Self-supervised SAR-optical Data Fusion and Land-cover Mapping using Sentinel-1/-2 Images
- Urban land-use analysis using proximate sensing imagery: a survey
- Mapping Vulnerable Populations with AI
- Cross-Modal Learning of Housing Quality in Amsterdam
- On What Depends the Robustness of Multi-source Models to Missing Data in Earth Observation?
- Learning a Dynamic Map of Visual Appearance