13 papers
RemoteZero: Geospatial Reasoning with Zero Labels
Liang Yao, Fan Liu, Shengxiang Xu +4
Geospatial reasoning requires models to identify image regions that satisfy complex and often implicit user intents. Recent reinforcement learning approaches improve reasoning with…
Evaluating Remote Sensing Image Captions Beyond Metric Biases
Ziyun Chen, Fan Liu, Liang Yao +3
The core objective of image captioning is to achieve lossless semantic compression from visual signals into textual modalities. However, the reliance on manually curated reference…
RemoteShield: Enable Robust Multimodal Large Language Models for Earth Observation
Rui Min, Liang Yao, Shiyu Miao +5
A robust Multimodal Large Language Model (MLLM) for Earth Observation should maintain consistent interpretation and reasoning under realistic input variations. However, current Rem…
RemoteAgent: Bridging Vague Human Intents and Earth Observation with RL-based Agentic MLLMs
Liang Yao, Shengxiang Xu, Fan Liu +7
Earth Observation (EO) systems are essentially designed to support domain experts who often express their requirements through vague natural language rather than precise, machine-f…
VisionNVS: Self-Supervised Inpainting for Novel View Synthesis under the Virtual-Shift Paradigm
Hongbo Lu, Liang Yao, Chenghao He +4
A fundamental bottleneck in Novel View Synthesis (NVS) for autonomous driving is the inherent supervision gap on novel trajectories: models are tasked with synthesizing unseen view…
Disentangle Object and Non-object Infrared Features via Language Guidance
Fan Liu, Ting Wu, Chuanyi Zhang +3
Infrared object detection focuses on identifying and locating objects in complex environments (\eg, dark, snow, and rain) where visible imaging cameras are disabled by poor illumin…