collaborators

10 papers

cs.AI2026

TouchThinker: Scaling Tactile Commonsense Reasoning to the Open World with Large-scale Data and Action-aware Representation

Kailin Lyu, Di Wu, Pengwei Zhang +12

Touch is a key modality for embodied agents to understand the physical world. Although recent work has incorporated tactile signals into language systems for tactile commonsense re…

cs.RO2026

BAT-Nav: Budget-Aware Arbitration and Termination for Long-Horizon Semantic Navigation

Xi Lin, Kangyi Wu, Jiayi Li +3

Long-horizon semantic navigation asks a robot to localize multiple open-vocabulary targets under a finite action budget. This setting exposes an execution failure that is largely h…

cs.CV2026

Dual-Anchoring: Addressing State Drift in Vision-Language Navigation

Kangyi Wu, Pengna Li, Kailin Lyu +5

Vision-Language Navigation(VLN) requires an agent to navigate through 3D environments by following natural language instructions. While recent Video Large Language Models(Video-LLM…

cs.CV2026

ATT-CR: Adaptive Triangular Transformer for Cloud Removal

Yang Wu, Ye Deng, Pengna Li +4

Cloud removal aims to accurately reconstruct the ground objects obscured by clouds in remote sensing images. Existing Transformer-based methods utilizing self-attention have shown…

cs.CV2026

SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation

Pengna Li, Kangyi Wu, Shaoqing Xu +7

Vision-and-Language Navigation (VLN) aims to enable an embodied agent to follow natural-language instructions and navigate to a target location in unseen 3D environments. We argue…

cs.RO2026

Think before Go: Hierarchical Reasoning for Image-goal Navigation

Pengna Li, Kangyi Wu, Shaoqing Xu +5

Image-goal navigation steers an agent to a target location specified by an image in unseen environments. Existing methods primarily handle this task by learning an end-to-end navig…