activity
20242026
collaborators

7 papers

cs.CV2026

RemoteZero: Geospatial Reasoning with Zero Labels

Liang Yao, Fan Liu, Shengxiang Xu +4

Geospatial reasoning requires models to identify image regions that satisfy complex and often implicit user intents. Recent reinforcement learning approaches improve reasoning with…

cs.CV2026

RemoteShield: Enable Robust Multimodal Large Language Models for Earth Observation

Rui Min, Liang Yao, Shiyu Miao +5

A robust Multimodal Large Language Model (MLLM) for Earth Observation should maintain consistent interpretation and reasoning under realistic input variations. However, current Rem…

cs.CV2026

RemoteAgent: Bridging Vague Human Intents and Earth Observation with RL-based Agentic MLLMs

Liang Yao, Shengxiang Xu, Fan Liu +7

Earth Observation (EO) systems are essentially designed to support domain experts who often express their requirements through vague natural language rather than precise, machine-f…

cs.CV2025

RemoteReasoner: Towards Unifying Geospatial Reasoning Workflow

Liang Yao, Fan Liu, Hongbo Lu +5

Remote sensing imagery presents vast, inherently unstructured spatial data, necessitating sophisticated reasoning to interpret complex user intents and contextual relationships bey…

cs.MA2025

RobustFlow: Towards Robust Agentic Workflow Generation

Shengxiang Xu, Jiayi Zhang, Shimin Di +6

The automated generation of agentic workflows is a promising frontier for enabling large language models (LLMs) to solve complex tasks. However, the empirical study reveals that ex…

cs.CV2025

UEMM-Air: Make Unmanned Aerial Vehicles Perform More Multi-modal Tasks

Liang Yao, Fan Liu, Shengxiang Xu +6

The development of multi-modal learning for Unmanned Aerial Vehicles (UAVs) typically relies on a large amount of pixel-aligned multi-modal image data. However, existing datasets f…