1 paper · 1 filter
Xinyue Xu, Zheng Zhang, Kunyang Ma +5
As vision-language models (VLMs) are increasingly deployed in geospatial question answering and visual scene understanding, improving their spatial cognition capability on street v…