5 papers
Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?
Qiwei Ma, Chunping Qiu, Xinjun Cheng +5
The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language…
Not All Patches are Equal: Sampling Matters for Visible-Infrared Pre-Training
Qiwei Ma, Bin Deng, Junjie Zhu +5
Visible-infrared (VIS-IR) alignment is a key pre-training task for robust multi-sensor perception. Most existing methods use uniform patch-wise contrastive learning, but this can b…
SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery
Qiwei Ma, Xukun Lu, Wang Liu +3
Synthetic Aperture Radar (SAR) is a critical imaging modality due to its all-weather operational capability. Although recent advances in self-supervised learning and masked image m…
DeepTaxon: An Interpretable Retrieval-Augmented Multimodal Framework for Unified Species Identification and Discovery
Jiawei Wang, Ming Lei, Yaning Yang +8
Identifying species in biology among tens of thousands of visually similar taxa while discovering unknown species in open-world environments remains a fundamental challenge in biod…
Differential Privacy Image Generation with Reconstruction Loss and Noise Injection Using an Error Feedback SGD
Qiwei Ma, Jun Zhang
Traditional data masking techniques such as anonymization cannot achieve the expected privacy protection while ensuring data utility for privacy-preserving machine learning. Synthe…