9 papers
CLOSER-VLN: Closed-Loop Self-Verified Retrieval-Augmented Reasoning for Aerial Vision-Language Navigation
Shaoxuan Li, Xiangyu Dong, Xiaoguang Ma +3
Vision-language navigation (VLN) has recently advanced with large language and multimodal models, enabling agents to follow natural-language instructions in unseen environments wit…
Acting While Understanding: Asynchronous Semantic-Action Decoupling for Real-Time Vision-Language-Action Models
Shenhao Yan, Ge Wang, Qi Liu +7
Vision-Language-Action models (VLAs) have demonstrated strong task understanding and generalization in robotic manipulation, yet the high computational cost of full-model inference…
LightZeroNav: Zero-Shot Vision Language Navigation in Continuous Environments Based on Lightweight VLMs
Kun Luo, Xiangyu Dong, Xiaoguang Ma +2
Although vision-language navigation (VLN) has progressed rapidly, zero-shot VLN in continuous environments (VLN-CE) remains highly challenging when using lightweight vision-languag…
DUPLEX: Agentic Dual-System Planning via LLM-Driven Information Extraction
Keru Hua, Ding Wang, Yaoying Gu +1
While Large Language Models (LLMs) provide semantic flexibility for robotic task planning, their susceptibility to hallucination and logical inconsistency limits their reliability…
Hierarchical Dual-Change Collaborative Learning for UAV Scene Change Captioning
Fuhai Chen, Pengpeng Huang, Junwen Wu +4
This paper proposes a novel task for UAV scene understanding - UAV Scene Change Captioning (UAV-SCC) - which aims to generate natural language descriptions of semantic changes in d…
PM-Nav: Priori-Map Guided Embodied Navigation in Functional Buildings
Jiang Gao, Xiangyu Dong, Haozhou Li +3
Existing language-driven embodied navigation paradigms face challenges in functional buildings (FBs) with highly similar features, as they lack the ability to effectively utilize p…