works on

From the 1 of 10 linked papers with an AI index.

activity
20242026
collaborators

10 papers

cs.CV2026

Think with Extra-Image: A Farmland Segmentation Agent Driven by Spatio-Temporal Information Gain

Haiyang Wu, Weiliang Mu, Zhuofei Du +4

The paper introduces FarmSeeker, a dynamic segmentation agent that detects ambiguous farmland regions in remote sensing images and queries additional spatio‑temporal data to improv…

cs.CV2026

iGSP:Implicit Gradient Subspace Projection for Efficient Continual Learning of Vision-Language Models

Xuezhi Cui, Dongbo Zhou, Wang Guo +8

Vision-Language Models require efficient adaptation to continually emerging downstream tasks. While Parameter-Efficient Fine-Tuning mitigates catastrophic forgetting, assigning iso…

cs.AI2026

RS-Claw: Progressive Active Tool Exploration via Hierarchical Skill Trees for Remote Sensing Agents

Liangtian Liu, Zeyuan Wang, Ziyu Li +8

The rise of multi-modal large language models (MLLMs) is shifting remote sensing (RS) intelligence from "see" to "action", as OpenClaw-style frameworks enable agents to autonomousl…

cs.AI2026

The Wittgensteinian Representation Hypothesis: Is Language the Attractor of Multimodal Convergence?

Zhaoyang Zhang, Run Shao, Dongyue Wu +4

Understanding why independently trained neural networks from different modalities converge toward shared representations, and where this convergence leads, remains an open question…

cs.AI2026

Remote Sensing Image Intelligent Interpretation with the Language-Centered Perspective: Principles, Methods and Challenges

Haifeng Li, Wang Guo, Haiyang Wu +6

The mainstream paradigm of remote sensing image interpretation has long been dominated by vision-centered models, which rely on visual features for semantic understanding. However,…

cs.CV2025

Towards Accurate UAV Image Perception: Guiding Vision-Language Models with Stronger Task Prompts

Mingning Guo, Mengwei Wu, Shaoxian Li +2

Existing image perception methods based on VLMs generally follow a paradigm wherein models extract and analyze image content based on user-provided textual task prompts. However, s…