collaborators

10 papers

cs.CV2026

SA-GEM: Scale-Adaptive and Geospatial Evidence-Modulated Token Pruning for Efficient Remote Sensing Large Vision-Language Models

Kexin Ma, Jing Xiao, Bowen Xing +2

RS-LVLMs have advanced multimodal understanding of Earth observation imagery, yet their performance is fundamentally constrained by high-resolution processing, as visual token coun…

quant-ph2026

Quantum-Inspired Phase Bicoherence Spectroscopy: A Framework for Detecting Universal Textural Angular Order Across Multi-Modal Complex Datasets

Zheng Xing, Chan-Tong Lam, Xiaochen Yuan

Classical image analysis routinely discards structurally meaningful orientation signatures encoded within Fourier phase, which are easily corrupted by local cellular rotation. Alth…

cs.CV2026

GeoMamba: A Geometry-driven MambaVision Framework and Dataset for Fine-grained Optical-SAR Object Retrieval

Tiantong Fang, Xiuwei Wang, Jing Xiao +3

Multi-source remote sensing enables complementary observation of ground objects, while cross-modal fine-grained object retrieval remains challenging, especially under unaligned opt…

cs.CV2026

DeTracker: Motion-decoupled Vehicle Detection and Tracking in Unstabilized Satellite Videos

Jiajun Chen, Jing Xiao, Shaohan Cao +4

Satellite videos provide continuous observations of surface dynamics but pose significant challenges for multi-object tracking (MOT), especially under unstabilized conditions where…

cs.CV2026

Observe Less, Understand More: Cost-aware Cross-scale Observation for Remote Sensing Understanding

Zhenghao Xie, Jing Xiao, Zhenqi Wang +5

Remote sensing understanding inherently requires multi-resolution observation, since different targets and application tasks demand different levels of spatial detail. While low-re…

cs.CV2026

Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models

Kexin Ma, Jing Xiao, Chaofeng Chen +4

Token pruning has emerged as an effective approach to reduce the substantial computational overhead of Large Vision-Language Models (LVLMs) by discarding less informative visual to…