9 papers
Geo-R1: Improving Few-Shot Geospatial Referring Expression Understanding with Reinforcement Fine-Tuning
Zilun Zhang, Zian Guan, Tiancheng Zhao +7
Referring expression understanding in remote sensing poses unique challenges, as it requires reasoning over complex object-context relationships. While supervised fine-tuning (SFT)…
Preserving Knowledge in Large Language Model with Model-Agnostic Self-Decompression
Zilun Zhang, Yutao Sun, Tiancheng Zhao +4
Humans can retain old knowledge while learning new information, but Large Language Models (LLMs) often suffer from catastrophic forgetting when post-pretrained or supervised fine-t…
Talking to Yourself: Defying Forgetting in Large Language Models
Yutao Sun, Mingshuai Chen, Tiancheng Zhao +5
Catastrophic forgetting remains a major challenge when fine-tuning large language models (LLMs) on narrow, task-specific data, often degrading their general knowledge and reasoning…
ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration
Haozhan Shen, Kangjia Zhao, Tiancheng Zhao +4
Multimodal Large Language Models (MLLMs) have demonstrated impressive capabilities in vision-language understanding. Recently, with the integration of test-time scaling techniques,…
ImageRAG: Enhancing Ultra High Resolution Remote Sensing Imagery Analysis with ImageRAG
Zilun Zhang, Haozhan Shen, Tiancheng Zhao +7
Ultra High Resolution (UHR) remote sensing imagery (RSI) (e.g. 100,000 100,000 pixels or more) poses a significant challenge for current Remote Sensing Multimodal Large La…
SRMF: A Data Augmentation and Multimodal Fusion Approach for Long-Tail UHR Satellite Image Segmentation
Yulong Guo, Zilun Zhang, Yongheng Shang +4
The long-tail problem presents a significant challenge to the advancement of semantic segmentation in ultra-high-resolution (UHR) satellite imagery. While previous efforts in UHR s…