3 papers
cs.CV2025
Towards Accurate UAV Image Perception: Guiding Vision-Language Models with Stronger Task Prompts
Mingning Guo, Mengwei Wu, Shaoxian Li +2
Existing image perception methods based on VLMs generally follow a paradigm wherein models extract and analyze image content based on user-provided textual task prompts. However, s…
cs.RO2025
BEDI: A Comprehensive Benchmark for Evaluating Embodied Agents on UAVs
Mingning Guo, Mengwei Wu, Jiarun He +3
With the rapid advancement of low-altitude remote sensing and Vision-Language Models (VLMs), Embodied Agents based on Unmanned Aerial Vehicles (UAVs) have shown significant potenti…
cs.CL2025
IFShip: Interpretable Fine-grained Ship Classification with Domain Knowledge-Enhanced Vision-Language Models
Mingning Guo, Mengwei Wu, Yuxiang Shen +2
End-to-end interpretation currently dominates the remote sensing fine-grained ship classification (RS-FGSC) task. However, the inference process remains uninterpretable, leading to…