7 papers
RemoteZero: Geospatial Reasoning with Zero Labels
Liang Yao, Fan Liu, Shengxiang Xu +4
Geospatial reasoning requires models to identify image regions that satisfy complex and often implicit user intents. Recent reinforcement learning approaches improve reasoning with…
RemoteShield: Enable Robust Multimodal Large Language Models for Earth Observation
Rui Min, Liang Yao, Shiyu Miao +5
A robust Multimodal Large Language Model (MLLM) for Earth Observation should maintain consistent interpretation and reasoning under realistic input variations. However, current Rem…
RemoteAgent: Bridging Vague Human Intents and Earth Observation with RL-based Agentic MLLMs
Liang Yao, Shengxiang Xu, Fan Liu +7
Earth Observation (EO) systems are essentially designed to support domain experts who often express their requirements through vague natural language rather than precise, machine-f…
RemoteReasoner: Towards Unifying Geospatial Reasoning Workflow
Liang Yao, Fan Liu, Hongbo Lu +5
Remote sensing imagery presents vast, inherently unstructured spatial data, necessitating sophisticated reasoning to interpret complex user intents and contextual relationships bey…
RobustFlow: Towards Robust Agentic Workflow Generation
Shengxiang Xu, Jiayi Zhang, Shimin Di +6
The automated generation of agentic workflows is a promising frontier for enabling large language models (LLMs) to solve complex tasks. However, the empirical study reveals that ex…
UEMM-Air: Make Unmanned Aerial Vehicles Perform More Multi-modal Tasks
Liang Yao, Fan Liu, Shengxiang Xu +6
The development of multi-modal learning for Unmanned Aerial Vehicles (UAVs) typically relies on a large amount of pixel-aligned multi-modal image data. However, existing datasets f…