7 papers · 1 filter
ATT-CR: Adaptive Triangular Transformer for Cloud Removal
Yang Wu, Ye Deng, Pengna Li +4
Cloud removal aims to accurately reconstruct the ground objects obscured by clouds in remote sensing images. Existing Transformer-based methods utilizing self-attention have shown…
SpaAct: Spatially-Activated Transition Learning with Curriculum Adaptation for Vision-Language Navigation
Pengna Li, Kangyi Wu, Shaoqing Xu +7
Vision-and-Language Navigation (VLN) aims to enable an embodied agent to follow natural-language instructions and navigate to a target location in unseen 3D environments. We argue…
Dual-Anchoring: Addressing State Drift in Vision-Language Navigation
Kangyi Wu, Pengna Li, Kailin Lyu +5
Vision-Language Navigation(VLN) requires an agent to navigate through 3D environments by following natural language instructions. While recent Video Large Language Models(Video-LLM…
HiMemVLN: Enhancing Reliability of Open-Source Zero-Shot Vision-and-Language Navigation with Hierarchical Memory System
Kailin Lyu, Kangyi Wu, Pengna Li +9
LLM-based agents have demonstrated impressive zero-shot performance in vision-language navigation (VLN) tasks. However, most zero-shot methods primarily rely on closed-source LLMs…
REGNav: Room Expert Guided Image-Goal Navigation
Pengna Li, Kangyi Wu, Jingwen Fu +1
Image-goal navigation aims to steer an agent towards the goal location specified by an image. Most prior methods tackle this task by learning a navigation policy, which extracts vi…
Camera-aware Label Refinement for Unsupervised Person Re-identification
Pengna Li, Kangyi Wu, Wenli Huang +2
Unsupervised person re-identification aims to retrieve images of a specified person without identity labels. Many recent unsupervised Re-ID approaches adopt clustering-based method…