5 papers
Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?
Qiwei Ma, Chunping Qiu, Xinjun Cheng +5
The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language…
Not All Patches are Equal: Sampling Matters for Visible-Infrared Pre-Training
Qiwei Ma, Bin Deng, Junjie Zhu +5
Visible-infrared (VIS-IR) alignment is a key pre-training task for robust multi-sensor perception. Most existing methods use uniform patch-wise contrastive learning, but this can b…
SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery
Qiwei Ma, Xukun Lu, Wang Liu +3
Synthetic Aperture Radar (SAR) is a critical imaging modality due to its all-weather operational capability. Although recent advances in self-supervised learning and masked image m…
NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment
Shuhao Han, Haotian Fan, Fangyuan Kong +112
This paper reports on the NTIRE 2025 challenge on Text to Image (T2I) generation model quality assessment, which will be held in conjunction with the New Trends in Image Restoratio…
Learning from Noisy Pseudo-labels for All-Weather Land Cover Mapping
Wang Liu, Zhiyu Wang, Xin Guo +3
Semantic segmentation of SAR images has garnered significant attention in remote sensing due to the immunity of SAR sensors to cloudy weather and light conditions. Nevertheless, SA…