1 paper
Xianzhi Ma, Shujun Wang, Xiaohan Li +3
Ultra-High-Resolution (UHR) remote sensing image understanding requires Vision-Language Models (VLMs) to capture both the global scene layout and sparse yet task-critical local det…