7 papers
GeoSolver: Scaling Test-Time Reasoning in Remote Sensing with Fine-Grained Process Supervision
Lang Sun, Ronghao Fu, Zhuoran Duan +3
While Vision-Language Models (VLMs) have significantly advanced remote sensing interpretation, enabling them to perform complex, step-by-step reasoning remains highly challenging.…
TravExplorer: Cross-Floor Embodied Exploration via Traversability-Aware 3-D Planning
Han Zheng, Zhe Chen, Yudong Huang +4
Zero-shot Object Navigation (ZSON) has shown promise for open-vocabulary target search in unseen environments, yet most existing systems remain tied to planar representations and s…
SkyNative: A Native Multimodal Framework for Remote Sensing Visual Evidence Reasoning
Xiao Yang, Ronghao Fu, Zhiwen Lin +10
Remote sensing vision-language models commonly rely on pretrained visual encoders to convert images into semantic features before language-model reasoning. While effective for scen…
Multi-Granularity Reasoning for Image Quality Assessment via Attribute-Aware Reinforcement Learning to Rank
Xiangyong Chen, Xiaochuan Lin, Haoran Liu +3
Recent advances in reasoning-induced image quality assessment (IQA) have demonstrated the power of reinforcement learning to rank (RL2R) for training vision-language models (VLMs)…
GeoDiT: A Diffusion-based Vision-Language Model for Geospatial Understanding
Jiaqi Liu, Ronghao Fu, Haoran Liu +2
Autoregressive models are structurally misaligned with the inherently parallel nature of geospatial understanding, forcing a rigid sequential narrative onto scenes and fundamentall…
OmniEarth: A Benchmark for Evaluating Vision-Language Models in Geospatial Tasks
Ronghao Fu, Haoran Liu, Weijie Zhang +4
Vision-Language Models (VLMs) have demonstrated effective perception and reasoning capabilities on general-domain tasks, leading to growing interest in their application to Earth o…