activity
20242026
collaborators

5 papers

cs.CV2026

Towards Vision-Language Geo-Foundation Model: A Survey

Yue Zhou, Zhihang Zhong, Xue Yang

Vision-Language Foundation Models (VLFMs) have made remarkable progress on various multimodal tasks, such as image captioning, image-text retrieval, visual question answering, and…

cs.CV2025

Multimodal Mathematical Reasoning Embedded in Aerial Vehicle Imagery: Benchmarking, Analysis, and Exploration

Yue Zhou, Litong Feng, Mengcheng Lan +5

Mathematical reasoning is critical for tasks such as precise distance and area computations, trajectory estimations, and spatial analysis in unmanned aerial vehicle (UAV) based rem…

cs.CV2025

GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding

Yue Zhou, Mengcheng Lan, Xiang Li +6

Remote sensing (RS) visual grounding aims to use natural language expression to locate specific objects (in the form of the bounding box or segmentation mask) in RS images, enhanci…

cs.CV2025

Text4Seg: Reimagining Image Segmentation as Text Generation

Mengcheng Lan, Chaofeng Chen, Yue Zhou +5

Multimodal Large Language Models (MLLMs) have shown exceptional capabilities in vision-language tasks; however, effectively integrating image segmentation into these models remains…

cs.CV2024

Revisiting the Integration of Convolution and Attention for Vision Backbone

Lei Zhu, Xinjiang Wang, Wayne Zhang +1

Convolutions (Convs) and multi-head self-attentions (MHSAs) are typically considered alternatives to each other for building vision backbones. Although some works try to integrate…