1 citations · 2 across the 15 of their papers we have counts for
22 papers · 1 filter
FIRM: Fine-Grained Intra-Token Representation of Masks for Remote Sensing Reasoning Segmentation
Weidong Tang, Kaiyu Li, Yikai Wang +4
Reasoning segmentation requires multimodal large language models (MLLMs) to translate implicit instructions into precise pixel-level masks. MLLMs encode an image as visual tokens,…
Driver2Map: Imitating Human Driving for Online High-Definition Map Construction
Pan Yin, Runtian Xia, Weisong Kuang +3
High-definition (HD) maps are essential for autonomous driving systems. In constructing such maps, onboard multi-view camera images, standard-definition maps and satellite images p…
OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation
Kaiyu Li, Zepeng Xin, Zixuan Jiang +4
Open-vocabulary Earth observation (EO) aims to localize geospatial concepts specified in natural language rather than a fixed label set. Existing benchmarks, however, usually cover…
CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming
Ruixun Liu, Lingyu Zhang, Lanxuan Xue +3
Humans can effortlessly reason about scenes across different viewpoints, yet it remains unclear whether Vision-Language Models (VLMs) possess similar cross-view spatial abilities.…
Multi-Modal Building Change Detection for Large-Scale Small Changes: Benchmark and Baseline
Ye Wang, Wei Lu, Zhihui You +6
Change detection in optical remote sensing imagery is susceptible to illumination fluctuations, seasonal changes, and variations in surface land-cover materials. Relying solely on…
Exchange Is All You Need for Remote Sensing Change Detection
Sijun Dong, Siming Fu, Kaiyu Li +3
Remote sensing change detection fundamentally relies on the effective fusion and discrimination of bi-temporal features. Prevailing paradigms typically utilize Siamese encoders bri…