activity
20242026
most citedSegEarth-R1: Geospatial Pixel Reasoning via Large Language Model

1 citations · 2 across the 15 of their papers we have counts for

collaborators
Showing cs.CVShow all

22 papers · 1 filter

cs.CV2026

FIRM: Fine-Grained Intra-Token Representation of Masks for Remote Sensing Reasoning Segmentation

Weidong Tang, Kaiyu Li, Yikai Wang +4

Reasoning segmentation requires multimodal large language models (MLLMs) to translate implicit instructions into precise pixel-level masks. MLLMs encode an image as visual tokens,…

cs.CV2026

Driver2Map: Imitating Human Driving for Online High-Definition Map Construction

Pan Yin, Runtian Xia, Weisong Kuang +3

High-definition (HD) maps are essential for autonomous driving systems. In constructing such maps, onboard multi-view camera images, standard-definition maps and satellite images p…

cs.CV2026

OVEarth-Bench: Evaluating Category Breadth and Query Diversity for Open-Vocabulary Earth Observation

Kaiyu Li, Zepeng Xin, Zixuan Jiang +4

Open-vocabulary Earth observation (EO) aims to localize geospatial concepts specified in natural language rather than a fixed label set. Existing benchmarks, however, usually cover…

cs.CV2026

CVSBench: A Comprehensive Benchmark for Cross-view Spatial Reasoning and Dreaming

Ruixun Liu, Lingyu Zhang, Lanxuan Xue +3

Humans can effortlessly reason about scenes across different viewpoints, yet it remains unclear whether Vision-Language Models (VLMs) possess similar cross-view spatial abilities.…

cs.CV2026

Multi-Modal Building Change Detection for Large-Scale Small Changes: Benchmark and Baseline

Ye Wang, Wei Lu, Zhihui You +6

Change detection in optical remote sensing imagery is susceptible to illumination fluctuations, seasonal changes, and variations in surface land-cover materials. Relying solely on…

cs.CV2026

Exchange Is All You Need for Remote Sensing Change Detection

Sijun Dong, Siming Fu, Kaiyu Li +3

Remote sensing change detection fundamentally relies on the effective fusion and discrimination of bi-temporal features. Prevailing paradigms typically utilize Siamese encoders bri…