9 papers
TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation
Kailin Lyu, Di Wu, Pengwei Zhang +12
Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense re…
Dual-Anchoring: Addressing State Drift in Vision-Language Navigation
Kangyi Wu, Pengna Li, Kailin Lyu +5
Vision-Language Navigation(VLN) requires an agent to navigate through 3D environments by following natural language instructions. While recent Video Large Language Models(Video-LLM…
ATT-CR: Adaptive Triangular Transformer for Cloud Removal
Yang Wu, Ye Deng, Pengna Li +4
Cloud removal aims to accurately reconstruct the ground objects obscured by clouds in remote sensing images. Existing Transformer-based methods utilizing self-attention have shown…
SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation
Pengna Li, Kangyi Wu, Shaoqing Xu +7
Vision-and-Language Navigation (VLN) aims to enable an embodied agent to follow natural-language instructions and navigate to a target location in unseen 3D environments. We argue…
Think before Go: Hierarchical Reasoning for Image-goal Navigation
Pengna Li, Kangyi Wu, Shaoqing Xu +5
Image-goal navigation steers an agent to a target location specified by an image in unseen environments. Existing methods primarily handle this task by learning an end-to-end navig…
HiMemVLN: Enhancing Reliability of Open-Source Zero-Shot Vision-and-Language Navigation with Hierarchical Memory System
Kailin Lyu, Kangyi Wu, Pengna Li +9
LLM-based agents have demonstrated impressive zero-shot performance in vision-language navigation (VLN) tasks. However, most zero-shot methods primarily rely on closed-source LLMs…