5 papers
DOFA-CLIP: Multimodal Vision-Language Foundation Models for Earth Observation
Zhitong Xiong, Yi Wang, Weikang Yu +7
Earth observation (EO) spans a broad spectrum of modalities, including optical, radar, multispectral, and hyperspectral data, each capturing distinct environmental signals. However…
On the Foundations of Earth and Climate Foundation Models
Xiao Xiang Zhu, Zhitong Xiong, Yi Wang +7
Foundation models have enormous potential in advancing Earth and climate sciences, however, current approaches may not be optimal as they focus on a few basic features of a desirab…
Vision-Language Models in Remote Sensing: Current Progress and Future Trends
Xiang Li, Congcong Wen, Yuan Hu +2
The remarkable achievements of ChatGPT and GPT-4 have sparked a wave of interest and research in the field of large language models for Artificial General Intelligence (AGI). These…
RRSIS: Referring Remote Sensing Image Segmentation
Zhenghang Yuan, Lichao Mou, Yuansheng Hua +1
Localizing desired objects from remote sensing images is of great use in practical applications. Referring image segmentation, which aims at segmenting out the objects to which a g…
ChatEarthNet: A Global-Scale Image-Text Dataset Empowering Vision-Language Geo-Foundation Models
Zhenghang Yuan, Zhitong Xiong, Lichao Mou +1
An in-depth comprehension of global land cover is essential in Earth observation, forming the foundation for a multitude of applications. Although remote sensing technology has adv…