From the 1 of 16 linked papers with an AI index.
16 papers
HAM-VLN: Harnessing Hierarchical Agentic Memory for Zero-Shot Vision-and-Language Navigation
An Liu, Bingxi Liu, Hongyu Ding +6
Vision-and-language navigation (VLN) enables robots to follow instructions in previously unseen environments. Recently, a training-free paradigm has emerged: the robot queries a mu…
ReferTrack: Referring Then Tracking for Embodied Visual Tracking
Hanjing Ye, Tianle Zeng, Jiazhao Zhang +6
Embodied visual tracking (EVT) requires a mobile agent to continuously follow a specific target described in natural language using only onboard vision. While recent vision-languag…
Traj-VLN: Learning Pixel-Space Interaction via Autoregressive Trajectory Generation
Changfei Fu, Guangcheng Chen, Wenjun Xu +3
The paper introduces a method that fine‑tunes vision‑language models to predict sequences of pixel coordinates, enabling an embodied agent to follow natural language navigation ins…
Can Single-View Mesh Reconstruction Generalize to Robot Camera Rotation?
Yu Zhan, Guangcheng Chen, Hanjing Ye +4
Single-view mesh reconstruction predicts object meshes and spatial layouts from a single observation, making it attractive for fast robot spatial reasoning and real-to-sim digital…
TARIC: Memory-Augmented Traversability-Aware Outdoor VLN under Interrupted Semantic Cues
Tianle Zeng, Hanjing Ye, Jianwei Peng +3
Outdoor vision-language navigation (VLN) in long-range, open-world environments is frequently disrupted by semantic-cue interruptions, where informative goal cues become sparse, oc…
Can Aerial VLA Models Cooperate? Evaluating Closed-Loop Air-Ground Coordination with CARLA-Air
Tianle Zeng, Yanci Wen, Xueang Yu +1
Recent aerial vision-language-action (VLA) models show promising single-UAV capabilities, such as tracking moving objects and navigating to language-specified landmarks. However, i…