2 papers
cs.CV2026
GRASP: Granularity-Aware Region Alignment and Semantic Prototype Learning for Fine-Grained Cross-Modal Understanding in Drone Views
Jiahui Cui, Yan Zhao, Kan Wei +4
Fine-grained cross-modal understanding in drone views is essential for aerial vision-language navigation. However, the inherent wide field of view and overhead perspective of drone…
cs.AI2026
Lingjing: A Simulation Testbed for Multi-Agent Embodied Tasks in Open-Ended Cities
Xiaohe Li, Yiru Wang, Junhao Fan +6
Urban embodied intelligence requires coordination among heterogeneous agents (e.g., UAVs, ground robots, and autonomous vehicles) in dynamic cities. Simulators therefore provide a…