5 papers
RSVideo: Are Your Vision-Language Models Ready for Remote Sensing Videos?
Hongjie Zhou, Shiqin Wang, Haoyang Chen +5
Remote-sensing videos enable real-time observation of changes in target attributes, short-term activities, and scene evolution. They record motion, actions, interactions, and scene…
PNEC-Mamba: Prototype-Guided Positive-Negative Evidence Calibration for Hyperspectral Image Classification
Mingzhen Xu, Can Xu, Di Wang +2
In real-world hyperspectral scenes, pixel representations are often ambiguous due to factors such as spectral similarity, mixed pixels, and local context interference, which may si…
MBTI: A Multi-Branch Efficient Fine-Tuning Framework for Hyperspectral Image Classification with Foundation Models
Mingzhen Xu, Haonan Guo, Di Wang +9
Hyperspectral foundation models learn transferable spectral-spatial representations from large-scale unlabeled data. They provide an effective paradigm for adapting to downstream h…
Seeing Clearly without Training: Mitigating Hallucinations in Multimodal LLMs for Remote Sensing
Yi Liu, Jing Zhang, Di Wang +3
Multimodal large language models (MLLMs) suffer from pronounced hallucinations in remote sensing visual question-answering (RS-VQA), primarily caused by visual grounding failures i…
SARMAE: Masked Autoencoder for SAR Representation Learning
Danxu Liu, Di Wang, Hebaixu Wang +6
Synthetic Aperture Radar (SAR) imagery plays a critical role in all-weather, day-and-night remote sensing applications. However, existing SAR-oriented deep learning is constrained…