5 papers
Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?
Qiwei Ma, Chunping Qiu, Xinjun Cheng +5
The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language…
Not All Patches are Equal: Sampling Matters for Visible-Infrared Pre-Training
Qiwei Ma, Bin Deng, Junjie Zhu +5
Visible-infrared (VIS-IR) alignment is a key pre-training task for robust multi-sensor perception. Most existing methods use uniform patch-wise contrastive learning, but this can b…
SARVLM: A Vision Language Foundation Model for Semantic Understanding in SAR Imagery
Qiwei Ma, Xukun Lu, Wang Liu +3
Synthetic Aperture Radar (SAR) is a critical imaging modality due to its all-weather operational capability. Although recent advances in self-supervised learning and masked image m…
Report of the 5th PVUW Challenge: Towards More Diverse Modalities in Pixel-Level Understanding
Chang Liu, Henghui Ding, Nikhila Ravi +40
This report summarizes the objectives, datasets, and top-performing methodologies of the 2026 Pixel-level Video Understanding in the Wild (PVUW) Challenge, hosted at CVPR 2026, whi…
MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
Peng Xu, Shengwu Xiong, Jiajun Zhang +125
This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We…