collaborators

5 papers

cs.CV2026

GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing

Aoran Xiao, Shihao Cheng, Yonghao Xu +3

Recent advances in multimodal large language models (MLLMs) have accelerated progress in domain-oriented AI, yet their development in geoscience and remote sensing (RS) remains con…

cs.CV2026

MM-OVSeg:Multimodal Optical-SAR Fusion for Open-Vocabulary Segmentation in Remote Sensing

Yimin Wei, Aoran Xiao, Hongruixuan Chen +2

Open-vocabulary segmentation enables pixel-level recognition from an open set of textual categories, allowing generalization beyond fixed classes. Despite great potential in remote…

cs.CV2025

A Vision Centric Remote Sensing Benchmark

Abduljaleel Adejumo, Faegheh Yeganli, Clifford Broni-bediako +3

Multimodal Large Language Models (MLLMs) have achieved remarkable success in vision-language tasks but their remote sensing (RS) counterpart are relatively under explored. Unlike n…

cs.CV2025

SARLANG-1M: A Benchmark for Vision-Language Modeling in SAR Image Understanding

Yimin Wei, Aoran Xiao, Yexian Ren +4

Synthetic Aperture Radar (SAR) is a crucial remote sensing technology, enabling all-weather, day-and-night observation with strong surface penetration for precise and continuous en…

cs.CV2025

OwlSight: A Robust Illumination Adaptation Framework for Dark Video Human Action Recognition

Shihao Cheng, Jinlu Zhang, Yue Liu +1

Human action recognition in low-light environments is crucial for various real-world applications. However, the existing approaches overlook the full utilization of brightness info…