2 papers
cs.CV2025
Towards Accurate UAV Image Perception: Guiding Vision-Language Models with Stronger Task Prompts
Mingning Guo, Mengwei Wu, Shaoxian Li +2
Existing image perception methods based on VLMs generally follow a paradigm wherein models extract and analyze image content based on user-provided textual task prompts. However, s…
cs.RO2025
BEDI: A Comprehensive Benchmark for Evaluating Embodied Agents on UAVs
Mingning Guo, Mengwei Wu, Jiarun He +3
With the rapid advancement of low-altitude remote sensing and Vision-Language Models (VLMs), Embodied Agents based on Unmanned Aerial Vehicles (UAVs) have shown significant potenti…