1 paper
Xiao Yang, Ronghao Fu, Zhiwen Lin +10
Remote sensing vision-language models (RS-VLMs) commonly employ a pretrained vision encoder and a projection module to map image features into the token space of a large language m…