collaborators

9 papers

cs.CV2026

CLOSER-VLN: Closed-Loop Self-Verified Retrieval-Augmented Reasoning for Aerial Vision-Language Navigation

Shaoxuan Li, Xiangyu Dong, Xiaoguang Ma +3

Vision-language navigation (VLN) has recently advanced with large language and multimodal models, enabling agents to follow natural-language instructions in unseen environments wit…

cs.RO2026

Acting While Understanding: Asynchronous Semantic-Action Decoupling for Real-Time Vision-Language-Action Models

Shenhao Yan, Ge Wang, Qi Liu +7

Vision-Language-Action models (VLAs) have demonstrated strong task understanding and generalization in robotic manipulation, yet the high computational cost of full-model inference…

cs.CV2026

LightZeroNav: Zero-Shot Vision Language Navigation in Continuous Environments Based on Lightweight VLMs

Kun Luo, Xiangyu Dong, Xiaoguang Ma +2

Although vision-language navigation (VLN) has progressed rapidly, zero-shot VLN in continuous environments (VLN-CE) remains highly challenging when using lightweight vision-languag…

cs.AI2026

DUPLEX: Agentic Dual-System Planning via LLM-Driven Information Extraction

Keru Hua, Ding Wang, Yaoying Gu +1

While Large Language Models (LLMs) provide semantic flexibility for robotic task planning, their susceptibility to hallucination and logical inconsistency limits their reliability…

cs.CV2026

Hierarchical Dual-Change Collaborative Learning for UAV Scene Change Captioning

Fuhai Chen, Pengpeng Huang, Junwen Wu +4

This paper proposes a novel task for UAV scene understanding - UAV Scene Change Captioning (UAV-SCC) - which aims to generate natural language descriptions of semantic changes in d…

cs.RO2026

PM-Nav: Priori-Map Guided Embodied Navigation in Functional Buildings

Jiang Gao, Xiangyu Dong, Haozhou Li +3

Existing language-driven embodied navigation paradigms face challenges in functional buildings (FBs) with highly similar features, as they lack the ability to effectively utilize p…