1 paper
Weizhi Chen, Yupeng Deng, Jin Wei +7
Vision Language Foundation Models based on CLIP architecture for remote sensing primarily rely on short text captions, which often result in incomplete semantic representations. Al…