7 papers
From Region Arrival to Instance-Level Grounding in Vision-and-Language Navigation
Xiangyu Shi, Ruoxi Yang, Wei Tao +3
Vision-and-Language Navigation (VLN) agents may satisfy conventional success criteria while still failing to establish reliable object-level grounding, because current evaluation p…
SpatialAnt: Autonomous Zero-Shot Robot Navigation via Active Scene Reconstruction and Visual Anticipation
Jiwen Zhang, Xiangyu Shi, Siyuan Wang +3
Vision-and-Language Navigation (VLN) has recently benefited from Multimodal Large Language Models (MLLMs), enabling zero-shot navigation. While recent exploration-based zero-shot m…
ROI-Driven Foveated Attention for Unified Egocentric Representations in Vision-Language-Action Systems
Xinhai Sun, Xiang Shi, Menglin Zou +1
The development of embodied AI systems is increasingly constrained by the availability and structure of physical interaction data. Despite recent advances in vision-language-action…
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
Jiwen Zhang, Zejun Li, Siyuan Wang +3
Although learning-based vision-and-language navigation (VLN) agents can learn spatial knowledge implicitly from large-scale training data, zero-shot VLN agents lack this process, r…
Fast-SmartWay: Panoramic-Free End-to-End Zero-Shot Vision-and-Language Navigation
Xiangyu Shi, Zerui Li, Yanyuan Qiao +1
Recent advances in Vision-and-Language Navigation in Continuous Environments (VLN-CE) have leveraged multimodal large language models (MLLMs) to achieve zero-shot navigation. Howev…
Embodied Domain Adaptation for Object Detection
Xiangyu Shi, Yanyuan Qiao, Lingqiao Liu +1
Mobile robots rely on object detectors for perception and object localization in indoor environments. However, standard closed-set methods struggle to handle the diverse objects an…