4 citations · 4 across the 1 of their papers we have counts for
3 papers
cs.CV2024
EarthMarker: A Visual Prompting Multi-modal Large Language Model for Remote Sensing
Wei Zhang, Miaoxin Cai, Tong Zhang +3
Recent advances in prompt learning have allowed users to interact with artificial intelligence (AI) tools in multi-turn dialogue, enabling an interactive understanding of images. H…
cs.CV2024
Popeye: A Unified Visual-Language Model for Multi-Source Ship Detection from Remote Sensing Imagery
Wei Zhang, Miaoxin Cai, Tong Zhang +3
Ship detection needs to identify ship locations from remote sensing (RS) scenes. Due to different imaging payloads, various appearances of ships, and complicated background interfe…
cs.CV2024★ 4 cited
EarthGPT: A Universal Multi-modal Large Language Model for Multi-sensor Image Comprehension in Remote Sensing Domain
Wei Zhang, Miaoxin Cai, Tong Zhang +2
Multi-modal large language models (MLLMs) have demonstrated remarkable success in vision and visual-language tasks within the natural image domain. Owing to the significant diversi…