2 papers
cs.RO2026
Vision-and-Language Navigation for UAVs: Progress, Challenges, and a Research Roadmap
Hanxuan Chen, Jie Zheng, Siqi Yang +9
Vision-and-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) represents a pivotal challenge in embodied artificial intelligence, focused on enabling UAVs to interpret high…
cs.CV2024
3DMIT: 3D Multi-modal Instruction Tuning for Scene Understanding
Zeju Li, Chao Zhang, Xiaoyan Wang +4
The remarkable potential of multi-modal large language models (MLLMs) in comprehending both vision and language information has been widely acknowledged. However, the scarcity of 3…