2 papers
cs.RO2026
Aerial World Model for Long-horizon Visual Generation and Navigation in 3D Space
Weichen Zhang, Peizhi Tang, Xin Zeng +12
Unmanned aerial vehicles (UAVs) have emerged as powerful embodied agents. One of the core abilities is autonomous navigation in large-scale three-dimensional environments. Existing…
cs.CV2025
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces
Baining Zhao, Jianjie Fang, Zichao Dai +8
Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban 3D space remain to be explored. We introduce a ben…